Jev vs Clef: two decision models, compared on what has been measured
Jev (TypeSafe) and Clef (Cloudflare) take the same request — a state plus typed Choice, Score and Noul questions — and return probabilities instead of text. Clef is open-weight, faster and self-serve; Jev is cheaper per token and has a much larger state budget — in our test on 2026-10-04, Clef ignored everything after roughly the first 2,000 tokens of the state. Use Clef-flash for short, high-volume decisions; use Jev when the input is long or the question needs judgment.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Route a ticket to the right team, rate its urgency and flag churn risk.
Edit freely — Jev answers the scenario's questions about whatever you put here.
At a glance
Clef is the first decision model from someone other than TypeSafe that copies Jev’s contract outright: Cloudflare launched it on 2026-10-01 as “fully Jev-API compatible”. That makes the comparison unusually clean, because you can send the same request to both. Every number below says where it comes from: (vendor) is the company selling the model, (our measurement) is a call we made ourselves.
| Jev | Clef | Clef-flash | |
|---|---|---|---|
| Made by | TypeSafe | Cloudflare | Cloudflare |
| Weights | Closed | Open, Apache 2.0 | Open, Apache 2.0 |
| Base model | Not disclosed | Qwen3.8-27B | Qwen3.5-9B |
| List price, input | $0.042 / M tokens | $0.24 / M tokens | $0.09 / M tokens |
| Output tokens | Free | Free | Free |
| Free hosted usage | None | 10,000 Neurons/day per account | 10,000 Neurons/day per account |
| Getting access | Waitlist | Self-serve Cloudflare account | Self-serve Cloudflare account |
| Questions per request | Up to 12 on JevStation | Up to 64 (vendor) | Up to 64 (vendor) |
| State actually read | Up to ~32k tokens shared with the longest question (vendor) | About the first 2,048 tokens (our measurement) | About the first 2,048 tokens (our measurement) |
| Images | No | Up to 4 per request | Up to 4 per request |
| Median latency, vendor | 524 ms | 209 ms | 39 ms |
| Typical evaluation, list price | $0.000020 (our measurement) | $0.000089 (our measurement) | $0.000033 (our measurement) |
| Credits on JevStation | 1 / 3 | 2 / 6 | 1 / 3 |
Credits are standard / large: an evaluation is large above 8,000 characters of state or five questions. The vendor latencies are Cloudflare’s numbers and are not comparable with each other — Clef was measured inside Cloudflare’s network and Jev across the public internet.
The difference that matters most: how much of the state is read
Cloudflare’s model page lists a 65,536-token context. In practice, on Workers AI, Clef reads far less of the state than that. We tested it on 2026-10-04 by hiding one decisive sentence — a customer asking to cancel — at different points in 20,000 characters of neutral text, and asking a single yes/no question about it:
| Where the sentence was | Clef-flash, p(cancel) | Clef, p(cancel) |
|---|---|---|
| Nowhere (control) | 0.02 | 0.01 |
| Start, 1,000, 6,000 or 12,000 characters in | 0.96–0.97 | 0.98–0.99 |
| At the very end | 0.02 | 0.01 |
Both models found the sentence anywhere in the first 12,000 characters and missed it entirely at the end. The token count Cloudflare billed stayed at 2,190 whatever we sent, which matches a cap of about 2,048 state tokens. Text past the cap is not read and not billed. The question set is not capped the same way: twelve long questions still counted as more than 8,000 tokens.
For English prose, 2,048 tokens is about 12,000 characters. For code, JSON or Chinese, Japanese and Korean text it is much less. Jev’s budget is about 32,000 tokens shared between the state and the longest question, so a long support thread, a contract or a full chat transcript fits. The full method and every number are in Cloudflare Clef, tested.
This may change — Cloudflare could lift the cap at any time — so check the date at the top of this page.
Benchmarks: what Cloudflare reports, and what nobody has checked
Cloudflare published results on its Jev Decision Index, 43 benchmarks across 10 metrics. These are vendor numbers, and at the time of writing no independent group has reproduced them.
| Task | Metric | Clef | Jev | Ahead |
|---|---|---|---|---|
| BANKING77 intent | macro-F1 | 94.20 | 79.74 | Clef |
| CLINC150 + out-of-scope | accuracy | 97.43 | 89.27 | Clef |
| ToolRet | nDCG@10 | 69.19 | 65.28 | Clef |
| BFCL tool call (Clef-flash) | exact case | 98.76 | 95.75 | Clef-flash |
| When2Call | accuracy | 72.37 | 80.97 | Jev |
| BRIGHT retrieval | nDCG@10 | 45.91 | 47.52 | Jev |
| Agent trace observability | accuracy | 68.5 | 71.6 | Jev |
Two cautions. First, do not mix these with older Jev numbers: the independent jev-benchmarks pilot measured Jev on Banking77 with plain accuracy on a 100-example sample, a different metric and a different sample from Cloudflare’s macro-F1. Second, the pattern is plausible on its face — Clef wins on fixed intent lists and tool lookup, Jev wins where the model has to decide whether to act at all — but that is a reading of vendor data, not a finding.
Price, honestly
Per token, Jev is the cheapest of the three. A typical evaluation — three questions and a short paragraph — cost us $0.000020 on Jev, $0.000033 on Clef-flash and $0.000089 on Clef at list price. At the other end, a request with 12 long questions costs up to $0.00038 on Jev and $0.0024 on Clef.
Cloudflare’s free allowance changes the picture for small volumes. Every account gets 10,000 Neurons a day, which at Clef’s rate covers about 458,000 input tokens: around 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day. If you already deploy on Cloudflare and are happy to write the integration, calling Workers AI yourself is the cheapest way to run Clef, and How to get Clef shows how.
JevStation charges per evaluation in credits, not per token. Clef-flash costs the same as Jev, 1 credit for a standard evaluation; Clef costs 2, because its token price is almost six times Jev’s. What you get for that is both models behind one key, a playground, batch runs, MCP tools, saved threads and automatic refunds when an evaluation fails. Pricing has a calculator for each model.
When to pick which
- The input is long (a thread, a document, a transcript): Jev. Clef will judge only the first part, silently.
- You need images judged: Clef or Clef-flash. Jev does not take images.
- Short inputs, high volume, latency matters: start with Clef-flash. It is the fastest of the three and costs the same number of credits as Jev here.
- Fixed intent lists, tool lookup: try Clef; Cloudflare’s own numbers favour it there.
- “Should the agent act at all?” questions: try Jev; it led on When2Call.
- The data cannot leave your machines: self-host Clef. Jev has no open weights.
- Lowest cost per token at scale: Jev, if you can get TypeSafe access.
Whatever the table says, the deciding test is your own data. Paste twenty real inputs into the playground, run them on both models, and compare the answers to what you know is right. For every other kind of alternative, including LLMs and trained classifiers, see Jev Alternatives and Jev vs open models.
Sources
- Launch, benchmarks, latency, data policy and fine-tuning. Cloudflare blog (vendor)
- Model IDs, limits and context window. Workers AI Clef model page and changelog (vendor)
- Prices and free Neurons. Workers AI pricing (vendor)
- Jev price and context. TypeSafe model docs (vendor)
- The 2,048-token plateau, first reported. Cloudflare community thread (independent)
- Truncation test, token counts and per-evaluation cost. Cloudflare Clef, tested and What One Jev Evaluation Costs (our measurement)
JevStation is independent. Jev and Clef are names of their respective owners, TypeSafe and Cloudflare; we are not affiliated with either.
Frequently asked questions
- Is Clef a drop-in replacement for Jev?
- Almost, for short inputs. Clef accepts the same state and the same Choice / Score / Noul questions and returns answers in the same shape. The two differences that break things: Workers AI reads only about the first 2,000 tokens of a Clef state (we measured this on 2026-10-04), and the REST API wraps the answer in a result object. Test your own question set before switching.
- Is Clef free?
- The weights are free under Apache 2.0. Hosted on Workers AI, Clef costs $0.24 and Clef-flash $0.09 per million input tokens, with output free, and every Cloudflare account gets 10,000 Neurons a day at no charge — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations. On JevStation, Clef-flash is in the free no-signup trial and both models run on the 200 sign-up credits.
- Does Clef generate text?
- No. Like Jev, it answers typed questions with probabilities and never writes prose, so it cannot replace an LLM for generation. It is a classifier and router you can describe in plain language.
- Can I self-host Clef?
- Yes. Both models are on Hugging Face under Apache 2.0; Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. Jev has no open weights and is only available as a hosted API.
- Which is more accurate, Jev or Clef?
- It depends on the task, and the only head-to-head so far is Cloudflare’s own. It reports Clef ahead on intent classification (BANKING77 macro-F1 94.2 against 79.7) and tool retrieval, and Jev ahead on deciding when to call a tool (When2Call 81.0 against 72.4), BRIGHT retrieval and agent-trace observability. No independent group has reproduced those numbers yet.
Related
- How to get ClefEvery way to use Cloudflare’s Clef: free trial, JevStation API, Workers AI direct, self-hosted — with prices and code.
- Jev vs open modelsopen-jev, jev-on-a-laptop, GLiNER and fine-tuned classifiers against hosted Jev: what each is, and what has been measured.
- How to get JevThe ways to access Jev — free trial, API key, TypeSafe direct — with prices, steps and code.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.