Jev vs Clef: two decision models, compared on what has been measured

Jev (TypeSafe) and Clef (Cloudflare) take the same request — a state plus typed Choice, Score and Noul questions — and return probabilities instead of text. Clef is open-weight, faster and self-serve; Jev is cheaper per token and has a much larger state budget — in our test on 2026-10-04, Clef ignored everything after roughly the first 2,000 tokens of the state. Use Clef-flash for short, high-volume decisions; use Jev when the input is long or the question needs judgment.

用你自己的文本试试

修改示例后运行。无需账号,每天免费 3 次。

场景

把工单分给对应团队,评估紧急程度,并标记流失风险。

166 / 2,000

可以随意修改,Jev 会针对这里的内容回答该场景的问题。

模型

At a glance

Clef is the first decision model from someone other than TypeSafe that copies Jev’s contract outright: Cloudflare launched it on 2026-10-01 as “fully Jev-API compatible”. That makes the comparison unusually clean, because you can send the same request to both. Every number below says where it comes from: (vendor) is the company selling the model, (our measurement) is a call we made ourselves.

JevClefClef-flash
Made byTypeSafeCloudflareCloudflare
WeightsClosedOpen, Apache 2.0Open, Apache 2.0
Base modelNot disclosedQwen3.8-27BQwen3.5-9B
List price, input$0.042 / M tokens$0.24 / M tokens$0.09 / M tokens
Output tokensFreeFreeFree
Free hosted usageNone10,000 Neurons/day per account10,000 Neurons/day per account
Getting accessWaitlistSelf-serve Cloudflare accountSelf-serve Cloudflare account
Questions per requestUp to 12 on JevStationUp to 64 (vendor)Up to 64 (vendor)
State actually readUp to ~32k tokens shared with the longest question (vendor)About the first 2,048 tokens (our measurement)About the first 2,048 tokens (our measurement)
ImagesNoUp to 4 per requestUp to 4 per request
Median latency, vendor524 ms209 ms39 ms
Typical evaluation, list price$0.000020 (our measurement)$0.000089 (our measurement)$0.000033 (our measurement)
Credits on JevStation1 / 32 / 61 / 3

Credits are standard / large: an evaluation is large above 8,000 characters of state or five questions. The vendor latencies are Cloudflare’s numbers and are not comparable with each other — Clef was measured inside Cloudflare’s network and Jev across the public internet.

The difference that matters most: how much of the state is read

Cloudflare’s model page lists a 65,536-token context. In practice, on Workers AI, Clef reads far less of the state than that. We tested it on 2026-10-04 by hiding one decisive sentence — a customer asking to cancel — at different points in 20,000 characters of neutral text, and asking a single yes/no question about it:

Where the sentence wasClef-flash, p(cancel)Clef, p(cancel)
Nowhere (control)0.020.01
Start, 1,000, 6,000 or 12,000 characters in0.96–0.970.98–0.99
At the very end0.020.01

Both models found the sentence anywhere in the first 12,000 characters and missed it entirely at the end. The token count Cloudflare billed stayed at 2,190 whatever we sent, which matches a cap of about 2,048 state tokens. Text past the cap is not read and not billed. The question set is not capped the same way: twelve long questions still counted as more than 8,000 tokens.

For English prose, 2,048 tokens is about 12,000 characters. For code, JSON or Chinese, Japanese and Korean text it is much less. Jev’s budget is about 32,000 tokens shared between the state and the longest question, so a long support thread, a contract or a full chat transcript fits. The full method and every number are in Cloudflare Clef, tested.

This may change — Cloudflare could lift the cap at any time — so check the date at the top of this page.

Benchmarks: what Cloudflare reports, and what nobody has checked

Cloudflare published results on its Jev Decision Index, 43 benchmarks across 10 metrics. These are vendor numbers, and at the time of writing no independent group has reproduced them.

TaskMetricClefJevAhead
BANKING77 intentmacro-F194.2079.74Clef
CLINC150 + out-of-scopeaccuracy97.4389.27Clef
ToolRetnDCG@1069.1965.28Clef
BFCL tool call (Clef-flash)exact case98.7695.75Clef-flash
When2Callaccuracy72.3780.97Jev
BRIGHT retrievalnDCG@1045.9147.52Jev
Agent trace observabilityaccuracy68.571.6Jev

Two cautions. First, do not mix these with older Jev numbers: the independent jev-benchmarks pilot measured Jev on Banking77 with plain accuracy on a 100-example sample, a different metric and a different sample from Cloudflare’s macro-F1. Second, the pattern is plausible on its face — Clef wins on fixed intent lists and tool lookup, Jev wins where the model has to decide whether to act at all — but that is a reading of vendor data, not a finding.

Price, honestly

Per token, Jev is the cheapest of the three. A typical evaluation — three questions and a short paragraph — cost us $0.000020 on Jev, $0.000033 on Clef-flash and $0.000089 on Clef at list price. At the other end, a request with 12 long questions costs up to $0.00038 on Jev and $0.0024 on Clef.

Cloudflare’s free allowance changes the picture for small volumes. Every account gets 10,000 Neurons a day, which at Clef’s rate covers about 458,000 input tokens: around 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day. If you already deploy on Cloudflare and are happy to write the integration, calling Workers AI yourself is the cheapest way to run Clef, and How to get Clef shows how.

JevStation charges per evaluation in credits, not per token. Clef-flash costs the same as Jev, 1 credit for a standard evaluation; Clef costs 2, because its token price is almost six times Jev’s. What you get for that is both models behind one key, a playground, batch runs, MCP tools, saved threads and automatic refunds when an evaluation fails. Pricing has a calculator for each model.

When to pick which

  • The input is long (a thread, a document, a transcript): Jev. Clef will judge only the first part, silently.
  • You need images judged: Clef or Clef-flash. Jev does not take images.
  • Short inputs, high volume, latency matters: start with Clef-flash. It is the fastest of the three and costs the same number of credits as Jev here.
  • Fixed intent lists, tool lookup: try Clef; Cloudflare’s own numbers favour it there.
  • “Should the agent act at all?” questions: try Jev; it led on When2Call.
  • The data cannot leave your machines: self-host Clef. Jev has no open weights.
  • Lowest cost per token at scale: Jev, if you can get TypeSafe access.

Whatever the table says, the deciding test is your own data. Paste twenty real inputs into the side-by-side runner, run them on both models at once, and compare the answers to what you know is right. For every other kind of alternative, including LLMs and trained classifiers, see Jev Alternatives and Jev vs open models.

Sources

JevStation is independent. Jev and Clef are names of their respective owners, TypeSafe and Cloudflare; we are not affiliated with either.

常见问题

Is Clef a drop-in replacement for Jev?
Almost, for short inputs. Clef accepts the same state and the same Choice / Score / Noul questions and returns answers in the same shape. The two differences that break things: Workers AI reads only about the first 2,000 tokens of a Clef state (we measured this on 2026-10-04), and the REST API wraps the answer in a result object. Test your own question set before switching.
Is Clef free?
The weights are free under Apache 2.0. Hosted on Workers AI, Clef costs $0.24 and Clef-flash $0.09 per million input tokens, with output free, and every Cloudflare account gets 10,000 Neurons a day at no charge — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations. On JevStation, Clef-flash is in the free no-signup trial and both models run on the 200 sign-up credits.
Does Clef generate text?
No. Like Jev, it answers typed questions with probabilities and never writes prose, so it cannot replace an LLM for generation. It is a classifier and router you can describe in plain language.
Can I self-host Clef?
Yes. Both models are on Hugging Face under Apache 2.0; Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. Jev has no open weights and is only available as a hosted API.
Which is more accurate, Jev or Clef?
It depends on the task, and the only head-to-head so far is Cloudflare’s own. It reports Clef ahead on intent classification (BANKING77 macro-F1 94.2 against 79.7) and tool retrieval, and Jev ahead on deciding when to call a tool (When2Call 81.0 against 72.4), BRIGHT retrieval and agent-trace observability. No independent group has reproduced those numbers yet.

相关页面

把它接进你的流程

注册即送 200 点,保存自己的问题集,并通过 API 调用同样的评估。