ブログに戻る

Cloudflare Clef Tested: Context Cap, Tokens, Cost

We called Clef and Clef-flash on Workers AI ~70 times: at launch they read only ~2,048 state tokens; a 10-07 retest found the cap gone. Costs per call.

更新日 JevStationJevStation

Clef is Cloudflare’s decision model: you send a state and a set of typed questions, and it returns a probability for each answer instead of text. Cloudflare released it on 2026-10-01 in two sizes, Clef (27B, built on Qwen3.8-27B) and Clef-flash (9B, built on Qwen3.5-9B), as open weights under Apache 2.0 and as hosted models on Workers AI (@cf/cloudflare/clef and @cf/cloudflare/clef-flash). It takes the same request as TypeSafe’s Jev, which is why we tested it.

Update, 2026-10-07: we re-ran the truncation test against Workers AI and the cap is gone. Both models found the sentence at the very end of 20,000 characters (p = 0.97 on Clef-flash, 0.99 on Clef), and the billed tokens grew with the text instead of stopping at 2,190. A 320,000-character state (48,062 tokens) was read to the end; a 640,000-character one was rejected with an error saying it exceeded the 65,536-token context window. Finding 1 below is what we measured on 2026-10-04 and is kept as a record.

Three days after launch we made about 70 calls to both models through the Workers AI REST API and recorded what came back. The short version:

  • At launch, only the first ~2,048 tokens of the state were read. Anything after that was ignored and not billed, although the model page says 65,536 tokens of context. Our 2026-10-07 retest found the whole state read.
  • A typical evaluation is 372 input tokens: $0.000089 on Clef and $0.000033 on Clef-flash at list price.
  • The question set is the expensive part. Twelve long questions cost 8,074 tokens before any state is added.
  • Clef-flash answered in 218 ms at the median from our machine, Clef in 507 ms. One Clef call in about 35 failed with an inference error.

What Clef is, in one paragraph

A decision model answers questions whose possible answers you list in advance. Clef supports the same three question types as Jev — choice (pick one named option), score (a level on an ordered scale) and noul (a yes/no probability) — up to 64 questions per request, plus up to four images. It does not generate text. Cloudflare says it does not read, store or train on requests and responses, except for customers of its fine-tuning service, which is currently limited to design partners.

How we tested

Every call used the REST endpoint POST /client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/{model} with a body of model, state and questions, and recorded usage.input_tokens from the response. The default question set was the three questions the JevStation playground starts with: one Choice with three options, one Noul and one five-level Score, 646 characters as JSON. For the worst cases we built questions with 600-character instructions and twelve 160-character options each, the most JevStation accepts. State text was neutral English filler.

Finding 1: on 2026-10-04, the state was cut at about 2,048 tokens

To see how much of the state Clef reads, we put one decisive sentence — a customer asking to cancel their subscription — inside 20,000 characters of neutral text, moved it around, and asked one Noul question: does the customer ask to cancel?

Where the sentence wasClef-flash, p(yes)Clef, p(yes)Tokens billed
Nowhere (control)0.0170.0102,190
First character0.9740.9872,190
1,000 characters in0.9580.9772,190
6,000 characters in0.9630.9842,190
12,000 characters in0.9740.9892,190
Last character (20,000)0.0170.0102,190

When the sentence was at the end, both models gave exactly the same answer as when it was absent. They did not see it. The billed count never moved from 2,190, which fits a fixed overhead plus a state budget of about 2,048 tokens — the same plateau a user reported on Cloudflare’s forum. The token table below shows it from the other side: 12,000 and 24,000 characters of state cost almost the same.

What 2,048 tokens holds depends on the text. In our English prose it was a little over 12,000 characters. Code, JSON and Chinese, Japanese or Korean text use more tokens per character, so less of them fits. There is no error or warning when the cut happens.

What we advised then. Put the part that decides the answer first, send the latest turns of a conversation rather than the oldest, or use a model with a larger budget. Since the 2026-10-07 retest found the cap gone, this no longer applies on Workers AI, and JevStation has dropped the long-state warning it showed for Clef.

Finding 2: the token model

Both models reported identical token counts for every input, so they share a tokenizer.

RequestQuestion set (chars)State (chars)Input tokens
1 Noul, 1-character state2051140
Default 3 questions, one sentence64658372
Default 3 questions6462,000698
Default 3 questions6468,0001,709
Default 3 questions64612,0002,378
Default 3 questions64624,0002,410
5 long questions13,7268,0004,728
12 long questions32,9431008,074
12 long questions32,94324,00010,106

Below the cap, the state costs about 0.168 tokens per character, roughly six characters per token. A typical question costs about 111 tokens, and long, fully described questions cost far more: the twelve-question set came to about 4.2 characters per token, because JSON keys, quotes and option lists tokenize densely. A rough fit for short question sets is:

input tokens ≈ 29 + 0.168 × min(state characters, ~12,000) + 111 × questions

It is an estimate for planning, not a bill. The real count is always in usage.input_tokens.

Compared with Jev, which we measured the same way, Clef has a much smaller fixed cost per request (Jev’s smallest call was 278 tokens against Clef’s 140) and spends more tokens on each question.

Finding 3: what an evaluation costs

At Cloudflare’s list prices — $0.24 per million input tokens for Clef, $0.09 for Clef-flash, output free:

RequestTokensClefClef-flashJev, for reference
Typical: 3 questions, a sentence372$0.000089$0.000033$0.000020
3 questions, 8,000 characters1,709$0.00041$0.00015—
5 long questions, 8,000 characters4,728$0.0011$0.00043—
Largest request JevStation accepts10,106$0.0024$0.00091$0.00038

The largest request was measured with the state cut at about 2,048 tokens. With the whole 24,000-character state read, as in our 2026-10-07 retest, it comes to about 11,960 tokens: $0.0029 on Clef and $0.0011 on Clef-flash.

A thousand typical evaluations cost about 9 cents on Clef and 3 cents on Clef-flash. Jev is cheaper per token ($0.042 per million), but Cloudflare gives every account 10,000 Neurons a day free. Clef is priced at 21,818 Neurons per million tokens, so the free allowance covers about 458,000 Clef tokens — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day.

On JevStation we charge per evaluation, in credits. Based on these numbers, Clef-flash costs the same as Jev (1 credit for a standard evaluation, 3 for a large one) and Clef costs twice as much (2 and 6). Pricing has a calculator for each model.

Finding 4: speed and reliability

We timed 15 identical requests per model (three questions, 400 characters) from a laptop over the public internet, so these numbers include the network round trip:

ModelMedian90th percentile
Clef-flash218 ms407 ms
Clef507 ms1,274 ms

Cloudflare reports 38.8 ms for Clef-flash and 209.3 ms for Clef, measured inside its own network. Ours are not comparable with those, only with each other: Clef-flash was about twice as fast and much more consistent.

In about 35 calls to Clef, one failed with HTTP 529 and error code 5012, “Clef inference failed”, on an ordinary 8,000-character request. Clef-flash returned no errors in a similar number of calls. Retry once on 5xx; JevStation does, and refunds the credits if the retry also fails.

The API, briefly

The request is the Jev request with a model field that must be clef or clef-flash. Through the REST API the answer is wrapped in Cloudflare’s envelope; through a Worker binding it is not:

{
  "result": {
    "model": "clef-flash",
    "answers": {
      "urgency": { "type": "noul", "noul": 0.929 }
    },
    "usage": { "input_tokens": 387, "output_tokens": 0 }
  },
  "success": true,
  "errors": [],
  "messages": []
}

How to get Clef has the full setup for Workers AI, for self-hosting and for JevStation.

Limits of this test

  • It is one afternoon of calls, three days after launch, from one location. Cloudflare can change the state cap, latency or prices at any time; the 2026-10-07 update at the top reflects our retest.
  • The truncation test used neutral English filler and one question. We did not test where exactly in a JSON or CJK state the cut lands.
  • We measured cost and behaviour, not accuracy. For accuracy, the only comparison so far is Cloudflare’s own, summarised with its caveats in Jev vs Clef.

Sources

  • State cap, token counts, latency and errors. Calls to Workers AI on 2026-10-04, truncation retested on 2026-10-07 (our measurement)
  • Model sizes, question types, limits and context window. Workers AI Clef model page and changelog (vendor)
  • Prices, Neurons per token and the free daily allowance. Workers AI pricing (vendor)
  • Base models, latency claims, data policy and fine-tuning. Cloudflare blog (vendor)
  • The 2,048-token plateau, first reported. Cloudflare community thread (independent)
  • Jev token counts. What One Jev Evaluation Costs (our measurement)

Try both models on your own input in the playground: new accounts get 200 free credits, no card.