Jev vs Clef: two decision models, compared on what has been measured
Jev (TypeSafe) and Clef (Cloudflare) take the same request — a state plus typed Choice, Score and Noul questions — and return probabilities instead of text. Clef is open-weight, faster and self-serve; Jev is cheaper per token. On 2026-10-04 Workers AI read only the first ~2,000 tokens of a Clef state; when we retested on 2026-10-07 that cap was gone. Use Clef-flash for short, high-volume decisions; use Jev when cost per token matters or the question needs judgment.
自分のテキストで試す
例を編集して実行してください。アカウントは不要で、1 日 3 回まで無料で実行できます。
シナリオ
チケットを適切なチームに振り分け、緊急度を評価し、解約リスクを検知します。
自由に編集できます。ここに入力した内容について、選んだモデルがシナリオの質問に回答します。
At a glance
Clef is the first decision model from someone other than TypeSafe that copies Jev’s contract outright: Cloudflare launched it on 2026-10-01 as “fully Jev-API compatible”. That makes the comparison unusually clean, because you can send the same request to both. Every number below says where it comes from: (vendor) is the company selling the model, (our measurement) is a call we made ourselves.
| Jev | Clef | Clef-flash | |
|---|---|---|---|
| Made by | TypeSafe | Cloudflare | Cloudflare |
| Weights | Closed | Open, Apache 2.0 | Open, Apache 2.0 |
| Base model | Not disclosed | Qwen3.8-27B | Qwen3.5-9B |
| List price, input | $0.042 / M tokens | $0.24 / M tokens | $0.09 / M tokens |
| Output tokens | Free | Free | Free |
| Free hosted usage | None | 10,000 Neurons/day per account | 10,000 Neurons/day per account |
| Getting access | Waitlist | Self-serve Cloudflare account | Self-serve Cloudflare account |
| Questions per request | Up to 12 on JevStation | Up to 64 (vendor) | Up to 64 (vendor) |
| State actually read | Up to ~32k tokens shared with the longest question (vendor) | All of it, 65,536-token context (our measurement, 2026-10-07) | All of it, 65,536-token context (our measurement, 2026-10-07) |
| Images | No | Up to 4 per request | Up to 4 per request |
| Median latency, vendor | 524 ms | 209 ms | 39 ms |
| Typical evaluation, list price | $0.000020 (our measurement) | $0.000089 (our measurement) | $0.000033 (our measurement) |
| Credits on JevStation | 1 / 3 | 2 / 6 | 1 / 3 |
Credits are standard / large: an evaluation is large above 8,000 characters of state or five questions. The vendor latencies are Cloudflare’s numbers and are not comparable with each other — Clef was measured inside Cloudflare’s network and Jev across the public internet.
How much of the state is read: a cap that has since gone
Cloudflare’s model page lists a 65,536-token context. When Clef launched, Workers AI read far less of the state than that. We tested it by hiding one decisive sentence — a customer asking to cancel — at different points in 20,000 characters of neutral text, and asking a single yes/no question about it. We ran the same test against Workers AI twice:
2026-10-04
| Where the sentence was | Clef-flash, p(cancel) | Clef, p(cancel) |
|---|---|---|
| Nowhere (control) | 0.02 | 0.01 |
| Start, 6,000 or 12,000 characters in | 0.96–0.97 | 0.98–0.99 |
| At the very end | 0.02 | 0.01 |
| Input tokens billed | 2,190 | 2,190 |
2026-10-07
| Where the sentence was | Clef-flash, p(cancel) | Clef, p(cancel) |
|---|---|---|
| Nowhere (control) | 0.02 | 0.00 |
| Start, 6,000 or 12,000 characters in | 0.95–0.97 | 0.99 |
| At the very end | 0.97 | 0.99 |
| Input tokens billed | 3,142–3,154 | 3,142–3,154 |
On 2026-10-04 both models missed the sentence at the end, and the billed count stayed at 2,190 whatever we sent: a cap of about 2,048 state tokens, with the rest neither read nor billed. On 2026-10-07 both found it at the end, and the billed count grew with the text. We kept going: a 320,000-character state (48,062 tokens) was read to the end, and a 640,000-character one was rejected with an error saying it exceeded the 65,536-token context window. So Workers AI now reads the whole state, and fails loudly rather than cutting silently when the input is too large.
Jev’s budget is about 32,000 tokens shared between the state and the longest question. Both are far more than the 24,000 characters a JevStation evaluation accepts. The original method and numbers are in Cloudflare Clef, tested.
Benchmarks: what Cloudflare reports, and what nobody has checked
Cloudflare published results on its Jev Decision Index, 43 benchmarks across 10 metrics. These are vendor numbers, and at the time of writing no independent group has reproduced them.
| Task | Metric | Clef | Jev | Ahead |
|---|---|---|---|---|
| BANKING77 intent | macro-F1 | 94.20 | 79.74 | Clef |
| CLINC150 + out-of-scope | accuracy | 97.43 | 89.27 | Clef |
| ToolRet | nDCG@10 | 69.19 | 65.28 | Clef |
| BFCL tool call (Clef-flash) | exact case | 98.76 | 95.75 | Clef-flash |
| When2Call | accuracy | 72.37 | 80.97 | Jev |
| BRIGHT retrieval | nDCG@10 | 45.91 | 47.52 | Jev |
| Agent trace observability | accuracy | 68.5 | 71.6 | Jev |
Two cautions. First, do not mix these with older Jev numbers: the independent jev-benchmarks pilot measured Jev on Banking77 with plain accuracy on a 100-example sample, a different metric and a different sample from Cloudflare’s macro-F1. Second, the pattern is plausible on its face — Clef wins on fixed intent lists and tool lookup, Jev wins where the model has to decide whether to act at all — but that is a reading of vendor data, not a finding.
Price, honestly
Per token, Jev is the cheapest of the three. A typical evaluation — three questions and a short paragraph — cost us $0.000020 on Jev, $0.000033 on Clef-flash and $0.000089 on Clef at list price. At the other end, a request with 12 long questions and 24,000 characters of state costs up to $0.00038 on Jev and about $0.0029 on Clef.
Cloudflare’s free allowance changes the picture for small volumes. Every account gets 10,000 Neurons a day, which at Clef’s rate covers about 458,000 input tokens: around 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day. If you already deploy on Cloudflare and are happy to write the integration, calling Workers AI yourself is the cheapest way to run Clef, and How to get Clef shows how.
JevStation charges per evaluation in credits, not per token. Clef-flash costs the same as Jev, 1 credit for a standard evaluation; Clef costs 2, because its token price is almost six times Jev’s. What you get for that is both models behind one key, a playground, batch runs, MCP tools, saved threads and automatic refunds when an evaluation fails. Pricing has a calculator for each model.
When to pick which
- The input is long (a thread, a document, a transcript): either reads it all now. Jev costs less per token; Clef has the larger context.
- You need images judged: Clef or Clef-flash. Jev does not take images.
- Short inputs, high volume, latency matters: start with Clef-flash. It is the fastest of the three and costs the same number of credits as Jev here.
- Fixed intent lists, tool lookup: try Clef; Cloudflare’s own numbers favour it there.
- “Should the agent act at all?” questions: try Jev; it led on When2Call.
- The data cannot leave your machines: self-host Clef. Jev has no open weights.
- Lowest cost per token at scale: Jev, if you can get TypeSafe access.
Whatever the table says, the deciding test is your own data. Paste twenty real inputs into the side-by-side runner, run them on both models at once, and compare the answers to what you know is right. For every other kind of alternative, including LLMs and trained classifiers, see Jev Alternatives and Jev vs open models.
Sources
- Launch, benchmarks, latency, data policy and fine-tuning. Cloudflare blog (vendor)
- Model IDs, limits and context window. Workers AI Clef model page and changelog (vendor)
- Prices and free Neurons. Workers AI pricing (vendor)
- Jev price and context. TypeSafe model docs (vendor)
- The 2,048-token plateau at launch, first reported. Cloudflare community thread (independent)
- Truncation tests (2026-10-04 and the 2026-10-07 retest), token counts and per-evaluation cost. Cloudflare Clef, tested and What One Jev Evaluation Costs (our measurement)
JevStation is independent. Jev and Clef are names of their respective owners, TypeSafe and Cloudflare; we are not affiliated with either.
よくある質問
- Is Clef a drop-in replacement for Jev?
- Almost. Clef accepts the same state and the same Choice / Score / Noul questions and returns answers in the same shape. The difference that breaks things is that the Workers AI REST API wraps the answer in a result object. On 2026-10-04 Workers AI also read only about the first 2,000 tokens of a Clef state; our retest on 2026-10-07 found the whole state read. Test your own question set before switching.
- Is Clef free?
- The weights are free under Apache 2.0. Hosted on Workers AI, Clef costs $0.24 and Clef-flash $0.09 per million input tokens, with output free, and every Cloudflare account gets 10,000 Neurons a day at no charge — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations. On JevStation, Clef-flash is in the free no-signup trial and both models run on the 200 sign-up credits.
- Does Clef generate text?
- No. Like Jev, it answers typed questions with probabilities and never writes prose, so it cannot replace an LLM for generation. It is a classifier and router you can describe in plain language.
- Can I self-host Clef?
- Yes. Both models are on Hugging Face under Apache 2.0; Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. Jev has no open weights and is only available as a hosted API.
- Which is more accurate, Jev or Clef?
- It depends on the task, and the only head-to-head so far is Cloudflare’s own. It reports Clef ahead on intent classification (BANKING77 macro-F1 94.2 against 79.7) and tool retrieval, and Jev ahead on deciding when to call a tool (When2Call 81.0 against 72.4), BRIGHT retrieval and agent-trace observability. No independent group has reproduced those numbers yet.
関連
- 意思決定モデル比較Jev・Clef・Clef-flash・Perplexity Decider・OpenAI Luna Decisions を同じ入力で実測:精度、キャリブレーション、コスト、レイテンシ。
- Clef playgroundTry Clef-flash free in the browser and compare Jev, Clef and Clef-flash side by side on your own input.
- How to get ClefEvery way to use Cloudflare’s Clef: free trial, JevStation API, Workers AI direct, self-hosted — with prices and code.
パイプラインに組み込む
新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。