Back to blog

Cloudflare Clef Tested: Context Cap, Tokens, Cost

We called Clef and Clef-flash on Workers AI ~70 times: they read only the first ~2,048 state tokens, a typical call is 372 tokens. Costs per call.

Updated JevStationJevStation

Clef is Cloudflare’s decision model: you send a state and a set of typed questions, and it returns a probability for each answer instead of text. Cloudflare released it on 2026-10-01 in two sizes, Clef (27B, built on Qwen3.8-27B) and Clef-flash (9B, built on Qwen3.5-9B), as open weights under Apache 2.0 and as hosted models on Workers AI (@cf/cloudflare/clef and @cf/cloudflare/clef-flash). It takes the same request as TypeSafe’s Jev, which is why we tested it.

Three days after launch we made about 70 calls to both models through the Workers AI REST API and recorded what came back. The short version:

  • Only the first ~2,048 tokens of the state are read. Anything after that is ignored and not billed. The model page says 65,536 tokens of context.
  • A typical evaluation is 372 input tokens: $0.000089 on Clef and $0.000033 on Clef-flash at list price.
  • The question set is the expensive part. Twelve long questions cost 8,074 tokens before any state is added.
  • Clef-flash answered in 218 ms at the median from our machine, Clef in 507 ms. One Clef call in about 35 failed with an inference error.

What Clef is, in one paragraph

A decision model answers questions whose possible answers you list in advance. Clef supports the same three question types as Jev — choice (pick one named option), score (a level on an ordered scale) and noul (a yes/no probability) — up to 64 questions per request, plus up to four images. It does not generate text. Cloudflare says it does not read, store or train on requests and responses, except for customers of its fine-tuning service, which is currently limited to design partners.

How we tested

Every call used the REST endpoint POST /client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/{model} with a body of model, state and questions, and recorded usage.input_tokens from the response. The default question set was the three questions the JevStation playground starts with: one Choice with three options, one Noul and one five-level Score, 646 characters as JSON. For the worst cases we built questions with 600-character instructions and twelve 160-character options each, the most JevStation accepts. State text was neutral English filler.

Finding 1: the state is cut at about 2,048 tokens

To see how much of the state Clef reads, we put one decisive sentence — a customer asking to cancel their subscription — inside 20,000 characters of neutral text, moved it around, and asked one Noul question: does the customer ask to cancel?

Where the sentence wasClef-flash, p(yes)Clef, p(yes)Tokens billed
Nowhere (control)0.0170.0102,190
First character0.9740.9872,190
1,000 characters in0.9580.9772,190
6,000 characters in0.9630.9842,190
12,000 characters in0.9740.9892,190
Last character (20,000)0.0170.0102,190

When the sentence was at the end, both models gave exactly the same answer as when it was absent. They did not see it. The billed count never moved from 2,190, which fits a fixed overhead plus a state budget of about 2,048 tokens — the same plateau a user reported on Cloudflare’s forum. The token table below shows it from the other side: 12,000 and 24,000 characters of state cost almost the same.

What 2,048 tokens holds depends on the text. In our English prose it was a little over 12,000 characters. Code, JSON and Chinese, Japanese or Korean text use more tokens per character, so less of them fits. There is no error or warning when the cut happens.

What to do about it. Put the part that decides the answer first. For a conversation, send the latest turns, not the oldest. For a long document, split it and ask per section, or use a model with a larger budget — Jev shares about 32,000 tokens between the state and its longest question. JevStation’s playground warns when a Clef state passes 8,000 characters.

Finding 2: the token model

Both models reported identical token counts for every input, so they share a tokenizer.

RequestQuestion set (chars)State (chars)Input tokens
1 Noul, 1-character state2051140
Default 3 questions, one sentence64658372
Default 3 questions6462,000698
Default 3 questions6468,0001,709
Default 3 questions64612,0002,378
Default 3 questions64624,0002,410
5 long questions13,7268,0004,728
12 long questions32,9431008,074
12 long questions32,94324,00010,106

Below the cap, the state costs about 0.168 tokens per character, roughly six characters per token. A typical question costs about 111 tokens, and long, fully described questions cost far more: the twelve-question set came to about 4.2 characters per token, because JSON keys, quotes and option lists tokenize densely. A rough fit for short question sets is:

input tokens ≈ 29 + 0.168 × min(state characters, ~12,000) + 111 × questions

It is an estimate for planning, not a bill. The real count is always in usage.input_tokens.

Compared with Jev, which we measured the same way, Clef has a much smaller fixed cost per request (Jev’s smallest call was 278 tokens against Clef’s 140) and spends more tokens on each question.

Finding 3: what an evaluation costs

At Cloudflare’s list prices — $0.24 per million input tokens for Clef, $0.09 for Clef-flash, output free:

RequestTokensClefClef-flashJev, for reference
Typical: 3 questions, a sentence372$0.000089$0.000033$0.000020
3 questions, 8,000 characters1,709$0.00041$0.00015—
5 long questions, 8,000 characters4,728$0.0011$0.00043—
Largest request JevStation accepts10,106$0.0024$0.00091$0.00038

A thousand typical evaluations cost about 9 cents on Clef and 3 cents on Clef-flash. Jev is cheaper per token ($0.042 per million), but Cloudflare gives every account 10,000 Neurons a day free. Clef is priced at 21,818 Neurons per million tokens, so the free allowance covers about 458,000 Clef tokens — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day.

On JevStation we charge per evaluation, in credits. Based on these numbers, Clef-flash costs the same as Jev (1 credit for a standard evaluation, 3 for a large one) and Clef costs twice as much (2 and 6). Pricing has a calculator for each model.

Finding 4: speed and reliability

We timed 15 identical requests per model (three questions, 400 characters) from a laptop over the public internet, so these numbers include the network round trip:

ModelMedian90th percentile
Clef-flash218 ms407 ms
Clef507 ms1,274 ms

Cloudflare reports 38.8 ms for Clef-flash and 209.3 ms for Clef, measured inside its own network. Ours are not comparable with those, only with each other: Clef-flash was about twice as fast and much more consistent.

In about 35 calls to Clef, one failed with HTTP 529 and error code 5012, “Clef inference failed”, on an ordinary 8,000-character request. Clef-flash returned no errors in a similar number of calls. Retry once on 5xx; JevStation does, and refunds the credits if the retry also fails.

The API, briefly

The request is the Jev request with a model field that must be clef or clef-flash. Through the REST API the answer is wrapped in Cloudflare’s envelope; through a Worker binding it is not:

{
  "result": {
    "model": "clef-flash",
    "answers": {
      "urgency": { "type": "noul", "noul": 0.929 }
    },
    "usage": { "input_tokens": 387, "output_tokens": 0 }
  },
  "success": true,
  "errors": [],
  "messages": []
}

How to get Clef has the full setup for Workers AI, for self-hosting and for JevStation.

Limits of this test

  • It is one afternoon of calls, three days after launch, from one location. Cloudflare can change the state cap, latency or prices at any time; we will update this post when we re-test.
  • The truncation test used neutral English filler and one question. We did not test where exactly in a JSON or CJK state the cut lands.
  • We measured cost and behaviour, not accuracy. For accuracy, the only comparison so far is Cloudflare’s own, summarised with its caveats in Jev vs Clef.

Sources

  • State cap, token counts, latency and errors. Calls to Workers AI on 2026-10-04 (our measurement)
  • Model sizes, question types, limits and context window. Workers AI Clef model page and changelog (vendor)
  • Prices, Neurons per token and the free daily allowance. Workers AI pricing (vendor)
  • Base models, latency claims, data policy and fine-tuning. Cloudflare blog (vendor)
  • The 2,048-token plateau, first reported. Cloudflare community thread (independent)
  • Jev token counts. What One Jev Evaluation Costs (our measurement)

Try both models on your own input in the playground: new accounts get 200 free credits, no card.