Jev vs Clef: two decision models, compared on what has been measured

Jev (TypeSafe) and Clef (Cloudflare) take the same request — a state plus typed Choice, Score and Noul questions — and return probabilities instead of text. Clef is open-weight, faster and self-serve; Jev is cheaper per token. On 2026-10-04 Workers AI read only the first ~2,000 tokens of a Clef state; when we retested on 2026-10-07 that cap was gone. Use Clef-flash for short, high-volume decisions; use Jev when cost per token matters or the question needs judgment.

自分のテキストで試す

例を編集して実行してください。アカウントは不要で、1 日 3 回まで無料で実行できます。

シナリオ

チケットを適切なチームに振り分け、緊急度を評価し、解約リスクを検知します。

166 / 2,000

自由に編集できます。ここに入力した内容について、選んだモデルがシナリオの質問に回答します。

モデル

At a glance

Clef is the first decision model from someone other than TypeSafe that copies Jev’s contract outright: Cloudflare launched it on 2026-10-01 as “fully Jev-API compatible”. That makes the comparison unusually clean, because you can send the same request to both. Every number below says where it comes from: (vendor) is the company selling the model, (our measurement) is a call we made ourselves.

JevClefClef-flash
Made byTypeSafeCloudflareCloudflare
WeightsClosedOpen, Apache 2.0Open, Apache 2.0
Base modelNot disclosedQwen3.8-27BQwen3.5-9B
List price, input$0.042 / M tokens$0.24 / M tokens$0.09 / M tokens
Output tokensFreeFreeFree
Free hosted usageNone10,000 Neurons/day per account10,000 Neurons/day per account
Getting accessWaitlistSelf-serve Cloudflare accountSelf-serve Cloudflare account
Questions per requestUp to 12 on JevStationUp to 64 (vendor)Up to 64 (vendor)
State actually readUp to ~32k tokens shared with the longest question (vendor)All of it, 65,536-token context (our measurement, 2026-10-07)All of it, 65,536-token context (our measurement, 2026-10-07)
ImagesNoUp to 4 per requestUp to 4 per request
Median latency, vendor524 ms209 ms39 ms
Typical evaluation, list price$0.000020 (our measurement)$0.000089 (our measurement)$0.000033 (our measurement)
Credits on JevStation1 / 32 / 61 / 3

Credits are standard / large: an evaluation is large above 8,000 characters of state or five questions. The vendor latencies are Cloudflare’s numbers and are not comparable with each other — Clef was measured inside Cloudflare’s network and Jev across the public internet.

How much of the state is read: a cap that has since gone

Cloudflare’s model page lists a 65,536-token context. When Clef launched, Workers AI read far less of the state than that. We tested it by hiding one decisive sentence — a customer asking to cancel — at different points in 20,000 characters of neutral text, and asking a single yes/no question about it. We ran the same test against Workers AI twice:

2026-10-04

Where the sentence wasClef-flash, p(cancel)Clef, p(cancel)
Nowhere (control)0.020.01
Start, 6,000 or 12,000 characters in0.96–0.970.98–0.99
At the very end0.020.01
Input tokens billed2,1902,190

2026-10-07

Where the sentence wasClef-flash, p(cancel)Clef, p(cancel)
Nowhere (control)0.020.00
Start, 6,000 or 12,000 characters in0.95–0.970.99
At the very end0.970.99
Input tokens billed3,142–3,1543,142–3,154

On 2026-10-04 both models missed the sentence at the end, and the billed count stayed at 2,190 whatever we sent: a cap of about 2,048 state tokens, with the rest neither read nor billed. On 2026-10-07 both found it at the end, and the billed count grew with the text. We kept going: a 320,000-character state (48,062 tokens) was read to the end, and a 640,000-character one was rejected with an error saying it exceeded the 65,536-token context window. So Workers AI now reads the whole state, and fails loudly rather than cutting silently when the input is too large.

Jev’s budget is about 32,000 tokens shared between the state and the longest question. Both are far more than the 24,000 characters a JevStation evaluation accepts. The original method and numbers are in Cloudflare Clef, tested.

Benchmarks: what Cloudflare reports, and what nobody has checked

Cloudflare published results on its Jev Decision Index, 43 benchmarks across 10 metrics. These are vendor numbers, and at the time of writing no independent group has reproduced them.

TaskMetricClefJevAhead
BANKING77 intentmacro-F194.2079.74Clef
CLINC150 + out-of-scopeaccuracy97.4389.27Clef
ToolRetnDCG@1069.1965.28Clef
BFCL tool call (Clef-flash)exact case98.7695.75Clef-flash
When2Callaccuracy72.3780.97Jev
BRIGHT retrievalnDCG@1045.9147.52Jev
Agent trace observabilityaccuracy68.571.6Jev

Two cautions. First, do not mix these with older Jev numbers: the independent jev-benchmarks pilot measured Jev on Banking77 with plain accuracy on a 100-example sample, a different metric and a different sample from Cloudflare’s macro-F1. Second, the pattern is plausible on its face — Clef wins on fixed intent lists and tool lookup, Jev wins where the model has to decide whether to act at all — but that is a reading of vendor data, not a finding.

Price, honestly

Per token, Jev is the cheapest of the three. A typical evaluation — three questions and a short paragraph — cost us $0.000020 on Jev, $0.000033 on Clef-flash and $0.000089 on Clef at list price. At the other end, a request with 12 long questions and 24,000 characters of state costs up to $0.00038 on Jev and about $0.0029 on Clef.

Cloudflare’s free allowance changes the picture for small volumes. Every account gets 10,000 Neurons a day, which at Clef’s rate covers about 458,000 input tokens: around 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations a day. If you already deploy on Cloudflare and are happy to write the integration, calling Workers AI yourself is the cheapest way to run Clef, and How to get Clef shows how.

JevStation charges per evaluation in credits, not per token. Clef-flash costs the same as Jev, 1 credit for a standard evaluation; Clef costs 2, because its token price is almost six times Jev’s. What you get for that is both models behind one key, a playground, batch runs, MCP tools, saved threads and automatic refunds when an evaluation fails. Pricing has a calculator for each model.

When to pick which

  • The input is long (a thread, a document, a transcript): either reads it all now. Jev costs less per token; Clef has the larger context.
  • You need images judged: Clef or Clef-flash. Jev does not take images.
  • Short inputs, high volume, latency matters: start with Clef-flash. It is the fastest of the three and costs the same number of credits as Jev here.
  • Fixed intent lists, tool lookup: try Clef; Cloudflare’s own numbers favour it there.
  • “Should the agent act at all?” questions: try Jev; it led on When2Call.
  • The data cannot leave your machines: self-host Clef. Jev has no open weights.
  • Lowest cost per token at scale: Jev, if you can get TypeSafe access.

Whatever the table says, the deciding test is your own data. Paste twenty real inputs into the side-by-side runner, run them on both models at once, and compare the answers to what you know is right. For every other kind of alternative, including LLMs and trained classifiers, see Jev Alternatives and Jev vs open models.

Sources

JevStation is independent. Jev and Clef are names of their respective owners, TypeSafe and Cloudflare; we are not affiliated with either.

よくある質問

Is Clef a drop-in replacement for Jev?
Almost. Clef accepts the same state and the same Choice / Score / Noul questions and returns answers in the same shape. The difference that breaks things is that the Workers AI REST API wraps the answer in a result object. On 2026-10-04 Workers AI also read only about the first 2,000 tokens of a Clef state; our retest on 2026-10-07 found the whole state read. Test your own question set before switching.
Is Clef free?
The weights are free under Apache 2.0. Hosted on Workers AI, Clef costs $0.24 and Clef-flash $0.09 per million input tokens, with output free, and every Cloudflare account gets 10,000 Neurons a day at no charge — roughly 1,200 typical Clef evaluations or 3,300 Clef-flash evaluations. On JevStation, Clef-flash is in the free no-signup trial and both models run on the 200 sign-up credits.
Does Clef generate text?
No. Like Jev, it answers typed questions with probabilities and never writes prose, so it cannot replace an LLM for generation. It is a classifier and router you can describe in plain language.
Can I self-host Clef?
Yes. Both models are on Hugging Face under Apache 2.0; Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B. Jev has no open weights and is only available as a hosted API.
Which is more accurate, Jev or Clef?
It depends on the task, and the only head-to-head so far is Cloudflare’s own. It reports Clef ahead on intent classification (BANKING77 macro-F1 94.2 against 79.7) and tool retrieval, and Jev ahead on deciding when to call a tool (When2Call 81.0 against 72.4), BRIGHT retrieval and agent-trace observability. No independent group has reproduced those numbers yet.

関連

パイプラインに組み込む

新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。