AI decision models compared: run OpenAI, Perplexity, TypeSafe and Cloudflare side by side
A decision model reads text or an image and answers questions you define in advance with probabilities, never prose. Since late September there are five you can call: TypeSafe’s Jev, Cloudflare’s Clef and Clef-flash, Perplexity’s Decider and OpenAI’s GPT-6 Luna Decisions. We ran all five on the same 196 labelled inputs on 2026-10-07. Clef-flash was the most accurate on fine-grained intents, Perplexity Decider the cheapest per single-question call, and no model won everything. Try four of them below without an account.
Try four decision models free
Pick OpenAI Luna Decisions, Perplexity Decider, Jev or Clef-flash, edit the example and run it. No account needed — 3 free runs a day.
Scenario
Route a ticket to the right team, rate its urgency and flag churn risk.
Edit freely — the model answers the scenario's questions about whatever you put here.
Run any two side by side
Same input, same questions, two models at once: Jev, Clef, Clef-flash, Perplexity Decider and OpenAI Luna Decisions. You see where the answers agree and what each one cost.
Scenario
What each model costs for your workload
Token counts come from our own calls on 2026-10-07, not from the vendors. Perplexity Decider reads the state once per question, so its cost grows with the number of questions.
JevTypeSafe · $0.042 / M input
- Direct, per month
- $2.66cheapest
- On JevStation, per month
- $137.10
- Input tokens / call
- 634
- Credits / call
- 1
Clef-flashCloudflare · $0.09 / M input
- Direct, per month
- $5.53
- On JevStation, per month
- $137.10
- Input tokens / call
- 614
- Credits / call
- 1
ClefCloudflare · $0.24 / M input
- Direct, per month
- $14.74
- On JevStation, per month
- $270.30
- Input tokens / call
- 614
- Credits / call
- 2
DeciderPerplexity · $0.04 / M input
- Direct, per month
- $3.94
- On JevStation, per month
- $270.30
- Input tokens / call
- 986
- Credits / call
- 2
Luna DecisionsOpenAI · $0.10 / M input
- Direct, per month
- $6.84
- On JevStation, per month
- $137.10
- Input tokens / call
- 684
- Credits / call
- 1
Estimates from token models fitted to our measured calls. Output tokens are free on all five. Direct prices exclude any gateway fee; the JevStation column is the cheapest mix of one-time credit packs after the 200 sign-up credits.
The five models at a glance
Every row says where its numbers come from: (vendor) is the company selling the model, (ours) is a call we made on 2026-10-07.
| Jev | Clef-flash | Clef | Perplexity Decider | OpenAI Luna Decisions | |
|---|---|---|---|---|---|
| Made by | TypeSafe | Cloudflare | Cloudflare | Perplexity | OpenAI |
| Launched | First of the five | 2026-10-01 | 2026-10-01 | 2026-10-02 | Previewed 2026-09-29 |
| Weights | Closed | Open, Apache 2.0 | Open, Apache 2.0 | Open, Apache 2.0 | Closed |
| Base model | Not disclosed | Qwen3.5-9B | Qwen3.8-27B | Qwen3.8-27B (vendor) | GPT-6 Luna |
| List price, input | $0.042 / M | $0.09 / M | $0.24 / M | $0.04 / M | $0.10 / M (OpenRouter) |
| Output tokens | Free | Free | Free | Free | Free |
| Context | ~32k shared (vendor) | 65k (vendor) | 65k (vendor) | 262k (vendor) | 1.05M (vendor) |
| Images | No | Yes | Yes | Yes | Yes |
| How the state is billed | Once | Once | Once | Once per question (ours) | Once |
| Median latency | 412 ms (ours) | 594 ms (ours) | 649 ms (ours) | 493 ms (ours) | 486 ms (ours) |
| Direct access | Waitlist | Self-serve | Self-serve | Self-serve | Limited preview; OpenRouter |
| Credits on JevStation | 1 / 3 | 1 / 3 | 2 / 6 | 2 / 6 | 1 / 3 |
| No-signup trial here | Yes | Yes | No | Yes | Yes |
Credits are standard / large: an evaluation is large above 8,000 characters of state or five questions. Our latencies are round trips from one machine to OpenRouter, so compare them with each other, not with vendor figures.
These are not the only decision models. OpenRouter listed 14 with a “decisions” output on the day we tested, including Liquid d1, Upstage Solar Decide, Inception Mercury Decide and Together’s experimental Tev1. We cover the five with the most behind them.
What we measured
Same inputs, same questions, same endpoint (OpenRouter’s /api/v1/systemone), five models, two full runs that agreed with each other. Every question was zero-shot: one line of plain description per option, nothing tuned.
- Banking77, card intents. 96 messages from the Banking77 test set across 12 intents that are easy to confuse: card arrival vs delivery estimate, lost vs compromised, declined vs not recognised. One 12-option Choice question.
- AG News. 100 headlines and leads from the AG News test set, 25 per topic. One four-option Choice question.
- Long state. One sentence asking to cancel a subscription, hidden at the start, 6,000 characters in, 12,000 characters in, or at the very end of 20,000 characters of neutral text, plus a control with no such sentence. One Noul question.
| Model | Banking77 cards (96) | AG News (100) | Banking77 answers at ≥ 0.9 confidence | Wrong at ≥ 0.99 confidence (of 196) |
|---|---|---|---|---|
| Jev | 86.5% | 91% | 81% of inputs, 94.9% right | 4 |
| Clef-flash | 97.9% | 92% | 73% of inputs, 100% right | 0 |
| Clef | 94.8% | 93% | 60% of inputs, 100% right | 0 |
| Perplexity Decider | 86.5% | 93% | 77% of inputs, 98.6% right | 1 |
| OpenAI Luna Decisions | 90.5% | 89% | 69% of inputs, 98.5% right | 6 |
What stands out:
- Clef-flash, the smallest model, was the most accurate on fine-grained intents. That matches Cloudflare’s own Banking77 claim, which until now nobody outside Cloudflare had checked.
- The Clef models are the most cautious. They rarely go above 0.99 and were never wrong when they did. Luna Decisions and Jev often answer with a confidence of 1.00, and that is where their mistakes cluster: six of Luna’s errors and four of Jev’s came at 0.99 or above. If you plan to auto-act above a threshold, set it on your own labelled data per model.
- OpenAI refused one input. The plain banking question “How do I freeze my account?” came back as an error, “OpenAI refused to answer”, in both runs. Plan for refusals as a failure mode; on JevStation the credit is refunded.
- All five read the whole 20,000-character state. Every model found the cancel request even at the very end. On 2026-10-04 we had measured Clef on Cloudflare Workers AI reading only about the first 2,048 tokens. On 2026-10-07 it read everything, both through OpenRouter and when we retested Workers AI directly, so that cap is gone.
With about 100 examples per task, a difference smaller than roughly six points could be noise. Only Clef-flash’s lead on Banking77 is clearly outside it. This is a small sample, not a leaderboard.
The cost catch: Perplexity Decider bills the state once per question
Decider has the lowest list price, and on a single question it was the cheapest model we ran: $0.007 per thousand AG News calls. But its token count grows with the number of questions, because it reads the whole state again for each one. On 8,000 characters of text, one question used 1,392 input tokens, three used 4,192 and five used 6,985. Luna Decisions, for comparison, used about 2,000 for five questions on the same text.
So the cheapest model depends on the shape of the request. Use the calculator above with your own numbers. JevStation charges Decider 2 credits for that reason; the other one-pass models cost 1.
How to choose
- Fixed intent lists, many similar labels: start with Clef-flash. It led our Banking77 run and costs 1 credit here.
- One question per call at high volume: Perplexity Decider, the cheapest per call in that shape.
- Several questions on long text: Jev or Luna Decisions, which read the state once.
- You want the OpenAI model without waiting for preview access: Luna Decisions through OpenRouter or here. Expect occasional refusals.
- The data cannot leave your machines: Clef, Clef-flash or Decider, all open-weight under Apache 2.0.
- You need a confidence you can threshold safely: test calibration on your own data; in our run the Clef models were the most conservative.
The deciding test is still your data. Paste twenty real inputs into the side-by-side runner above, run two models at once and compare the answers with what you know is right.
More comparisons
- Jev vs Perplexity Decider: the two cheapest models, head to head.
- OpenAI Decisions API alternatives: what to use while the preview is closed, and how to port a request.
- Jev vs Clef: Cloudflare’s models, including the Workers AI truncation we measured at launch and the retest that found it gone.
- Decision model pricing calculator: monthly cost for all five on your workload.
Sources
- OpenAI Decisions API preview, DevDay 2026-09-29. Coverage on Hugging Face and r/OpenAI (independent)
- Perplexity Decider model, price and context. OpenRouter model page and Perplexity Decisions API docs (vendor)
- System One endpoint and request format. OpenRouter TypeSafe SDK guide (vendor)
- Clef. Cloudflare blog (vendor) and Cloudflare Clef, tested (ours)
- Datasets. Banking77 and AG News test splits
- Accuracy, calibration, latency, token counts and long-state test. Our runs on 2026-10-07 (ours)
JevStation is independent. Jev, Clef, Decider and GPT-6 Luna are names of their respective owners, TypeSafe, Cloudflare, Perplexity and OpenAI; we are not affiliated with any of them.
Frequently asked questions
- What is a decision model?
- A model that returns a typed answer instead of text: a choice from options you list, a level on a rubric you define, or the probability that a statement is true. You send the content as a state and the questions as a schema, and you get back probabilities your code can threshold. TypeSafe calls the category System One; OpenAI calls its version the Decisions API.
- Which decision model is the most accurate?
- It depends on the task. In our run on 2026-10-07, Clef-flash scored 97.9% on 12 confusable Banking77 card intents, ahead of Clef (94.8%), OpenAI Luna Decisions (90.5%), Jev and Perplexity Decider (86.5% each). On four-way AG News topics all five were within 89–93%. With about 100 examples per task, differences under roughly 6 points are within noise.
- Can I use the OpenAI Decisions API without the waitlist?
- OpenAI’s own Decisions API was still in limited preview when we checked. OpenRouter has served the same model as openai/gpt-6-luna-decisions since 2026-10-06, and JevStation runs it through OpenRouter: free in the no-signup trial, 1 credit per standard evaluation with an account.
- Which decision model is the cheapest?
- Per token, Perplexity Decider at $0.04 per million input tokens, then Jev at $0.042. But Decider reads the state once for every question, so a five-question call on an 8,000-character text used 6,985 tokens in our test against about 2,000 for Luna Decisions. For one question Decider is the cheapest; for many questions on long text it is not.
- Are the request formats compatible?
- Yes. All five accept the same state plus Choice, Score and Noul questions and return answers in the same shape when called through OpenRouter’s System One endpoint or JevStation. Switching models is a change to one model field.
- Do decision models replace LLMs?
- No. They cannot write, summarise or reason in prose. They replace the LLM call you use as a classifier, router, judge or guardrail, where the possible answers can be listed in advance.
Related
- Decisions API alternativesOpenAI Decisions API without the waitlist, plus Jev, Clef and Perplexity Decider: measured accuracy, cost, and how to port a request.
- Jev vs DeciderPerplexity Decider against Jev: same contract, measured accuracy, calibration, latency, and per-question billing.
- Model pricingMonthly cost of five decision models for your volume, text length and question count, from measured token counts.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.