Back to blog

What One Jev Evaluation Costs, Measured

We made 11 real Jev calls across every state size and question count the app allows, fitted the token model, and measured the cost: $0.000012 to $0.000377 per evaluation.

Updated JevStationJevStation
What One Jev Evaluation Costs, Measured

A Jev evaluation costs between $0.000012 and $0.000377, depending on how much state you send and how many questions you ask. The cheap end is a one-line ticket with a single Noul question; the expensive end is the largest request JevStation will accept — a 24,000-character state with twelve fully-specified questions.

That is a 31x spread, and almost none of it comes from the part you would guess. We measured it rather than estimated it, because the credit tiers have to be safe at the ceiling, not the average.

Why measure instead of estimate

TypeSafe publishes one number: $0.042 per million input tokens, output free. That is enough to price a token, and not enough to price an evaluation. Two questions decide the real cost, and the vendor's page does not answer either:

  1. How many tokens does a request actually use? The state and the question set are different kinds of text — one is prose, one is JSON — and they do not tokenise at the same rate.
  2. Is there a fixed cost per request? If every call carries overhead, small evaluations are dominated by it and a per-token model will underprice them.

So we sent 11 real requests to the live API, spanning state sizes from 1 character to 24,000 and question counts from 1 to 12, and recorded the token usage Jev reported back.

The token model

Input tokens fit a straight line with an intercept:

input_tokens ≈ 260 + 0.163 × state_chars + 0.20 × question_chars

Read as rates, that is 260 tokens of fixed overhead per request, then roughly 6 characters per token for the state and 5 characters per token for the question set.

ComponentRateWhy
Fixed overhead260 tokens ($0.000011)TypeSafe's own request envelope, paid on every call
state~6 characters / tokenOrdinary English prose
questions~5 characters / tokenJSON is punctuation-heavy: keys, type, quotes and commas all cost

The two findings worth keeping:

  • Fixed overhead is more than half of a minimal call. 260 tokens is $0.0000109, and the smallest request we could construct cost $0.000012. At the bottom of the range you are paying for the envelope, not your content.
  • A question set costs more per character than the state it judges — 0.20 tokens per character against 0.163, so about 23% more. JSON structure is not free. Twelve questions with full criteria serialise to 24,143 characters and consume roughly 4,800 tokens before your state is counted at all.

Fitted against measured

The model is a fit, so here it is against the calls it was fitted to. The middle rows are the ones that matter — they are the shapes real usage takes.

Callstate charsQuestion charsMeasured tokensFittedCost
1 Noul question, 1-character state190278278$0.000012
Default 3 questions, real playground call58646484399$0.000020
Default 3 questions, 2,000-character state2,000646773715$0.000032
12 full questions, 100-character state10024,1435,0945,105$0.000214
Application ceiling24,00024,1438,9819,001$0.000377

One deviation is worth flagging rather than hiding: the fit underestimates small question sets by 10–18%. The default question set carries criteria written as compact structured strings, which tokenise worse than prose of the same length — so the linear coefficient is slightly low at that end. It is the reason the numbers below are taken from the measured column and never from the fitted one.

What each shape costs

ScenarioInput tokensCost
Minimum call (1 question, one sentence)278$0.000012
Playground quick check (3 questions, a paragraph)484$0.000020
Playground, real document (3 questions, 5k chars)~1,200$0.000050
Playground, long input (3 questions, 24k chars)~4,300$0.000180
API, lean (1 question, 1k state)441$0.000019
API, heavy (12 questions, 8k state)6,378$0.000268
Absolute ceiling (12 questions + 24k state)8,981$0.000377

The anchor to remember: 1,000 evaluations cost about 2 cents at the typical shape, or 38 cents if every one of them is the maximum-size request.

How this maps to credits

JevStation charges in credits rather than tokens, because a caller should not have to model tokenisation to know what a run costs. The tiers come straight from the two thresholds above:

  • A standard evaluation costs 1 credit — state up to 8,000 characters and up to 5 questions.
  • A large evaluation costs 3 credits — state up to 24,000 characters and up to 12 questions.
  • A failed evaluation costs 0 credits. Credits are consumed when the answers land and refunded automatically if the provider errors.

Because the ceiling is measured rather than fitted, the 3-credit tier is safe for every request the application will accept.

Reproduce it

The measurement is a script in the repository, not a table typed into a post:

# 8 calibration calls — fits the token model
pnpm verify:jev-cost

# 3 calls at the application ceiling — measures the cost ceiling
pnpm verify:jev-worst

# Recompute the cost tables from the fitted model
pnpm pricing:model

Eleven calls in total, about $0.0004 of real API usage.

Limits of this measurement

Stated plainly, because a benchmark without its boundaries is marketing:

  • One provider, one model version. Every call went to jev-latest, which resolved to a single build. Aliases move. If the tokeniser or the request envelope changes, these coefficients are stale.
  • Eleven calls. That is enough to fit two variables and an intercept, and not enough to characterise rare inputs. Non-English state, heavy Unicode, or unusual JSON shapes may tokenise differently — we did not measure them.
  • Vendor pricing is an input, not a constant. The $0.042 per million figure and the free output tokens are TypeSafe's published terms and can change; the token counts would survive that, the dollar figures would not.
  • This is model cost only. It is what one evaluation costs to run, and says nothing about what it should cost you — that is a product and pricing question, not a measurement.

If you re-run the scripts against a different build and get different coefficients, that is a useful result and we would like to hear about it.

To see what your own inputs cost, paste them into the playground — the first 200 evaluations are free, and the pricing page lists what each credit pack works out to per 1,000 evaluations.