Jev vs Perplexity Decider: the two cheapest decision models, tested on the same inputs
Perplexity Decider (pplx-decider-v1-27b) and Jev take the same request, a state plus typed questions, and return probabilities. They are also the two cheapest per token: $0.04 and $0.042 per million input tokens. In our run on 2026-10-07 they tied on Banking77 card intents (86.5%) and Decider edged ahead on AG News (93% against 91%). The difference that matters is billing: Decider reads the state again for every question, so it is the cheaper model for one question and the more expensive one for several on long text.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Route a ticket to the right team, rate its urgency and flag churn risk.
Edit freely — the model answers the scenario's questions about whatever you put here.
At a glance
(vendor) is the company selling the model; (ours) is a call we made on 2026-10-07 through OpenRouter.
| Jev | Perplexity Decider | |
|---|---|---|
| Made by | TypeSafe | Perplexity |
| Model id | jev-latest | pplx-decider-v1-27b |
| Weights | Closed | Open, Apache 2.0 |
| Base model | Not disclosed | Qwen3.8-27B (vendor) |
| List price, input | $0.042 / M tokens | $0.04 / M tokens |
| Output tokens | Free | Free |
| Context | ~32k tokens, shared with the longest question (vendor) | 262k tokens (vendor) |
| Images | No | Yes |
| Input tokens, our fit | ≈ 260 + 0.163 × characters + 43 × questions | ≈ questions × (84 + 0.163 × characters) |
| Median latency | 412 ms (ours) | 493 ms (ours) |
| Banking77 card intents | 86.5% (ours) | 86.5% (ours) |
| AG News topics | 91% (ours) | 93% (ours) |
| Credits on JevStation | 1 / 3 | 2 / 6 |
Credits are standard / large; an evaluation is large above 8,000 characters of state or five questions.
Accuracy and calibration
We sent both models the same 196 labelled inputs, zero-shot, with one line of description per option. The method is on Decision models compared, which runs all five models on the same data.
| Task | Jev | Decider |
|---|---|---|
| Banking77, 12 card intents (96) | 86.5% | 86.5% |
| AG News, 4 topics (100) | 91% | 93% |
| Banking77: answers at ≥ 0.9 confidence | 81% of inputs, 94.9% right | 77% of inputs, 98.6% right |
| Wrong at ≥ 0.99 confidence (of 196) | 4 | 1 |
| Expected calibration error, AG News | 0.086 | 0.052 |
The accuracy is a tie at this sample size. Calibration is where Decider pulled ahead: when it was very sure, it was nearly always right, so a threshold set on its confidence lets fewer mistakes through. Jev gave a confidence of 1.00 (to two decimals) on 136 of the 196 inputs, which leaves little room to separate easy cases from risky ones.
Both models found a single cancel request hidden at the very end of 20,000 characters of neutral text, so neither truncated the state in our test.
The cost difference is in the token count, not the price
The list prices are almost the same. The bills are not, because the two models count tokens differently. We fitted both on nine requests each (state of 100, 2,000 and 8,000 characters, with one, three and five questions):
| Request | Jev tokens | Decider tokens | Cheaper at list price |
|---|---|---|---|
| 1 question, short ticket (~300 chars) | ~350 | ~130 | Decider |
| 3 questions, 1,500 characters | ~640 | ~990 | Jev |
| 5 questions, 8,000 characters | ~1,780 | 6,985 | Jev, by about 4× |
| 12 long questions, 24,000 characters | 8,981 | 51,882 | Jev, by about 6× |
Decider’s count grows with questions × state length, as if each question were a separate call. Jev’s grows with state length plus a small amount per question. If your workload is one question per call, such as a router or a single guardrail, Decider is the cheapest model we have measured. If you ask several things about the same document, Jev is cheaper, even though its list price is higher.
On JevStation that is why Decider costs 2 credits rather than 1. The pricing calculator shows both for your own numbers.
When to pick which
- One question per call at high volume: Decider.
- Several questions on the same text: Jev.
- You need images judged: Decider. Jev does not take images.
- You need a confidence you can act on automatically: Decider was better calibrated in our run; confirm on your own data.
- Self-hosting or data residency: Decider, open-weight. Jev is API-only.
- Inputs longer than ~30k tokens: Decider, with a 262k context.
Try both
Run Decider above without an account, or sign in and use the side-by-side runner to send the same text and questions to Jev and Decider at once.
Sources
- Decider model, price and context. OpenRouter model page and Perplexity Decisions API docs (vendor)
- Launch, base model and licence. AGTP on X and AGI Hunt (independent)
- Jev price and context. TypeSafe model docs (vendor); worst-case Jev token count from What One Jev Evaluation Costs (ours)
- Accuracy, calibration, latency and token counts. Our runs on 2026-10-07 through OpenRouter (ours)
JevStation is independent. Jev and Decider are names of their respective owners, TypeSafe and Perplexity; we are not affiliated with either.
Frequently asked questions
- Is Perplexity Decider a drop-in replacement for Jev?
- For the request format, yes. Through OpenRouter’s System One endpoint or JevStation, Decider accepts the same state and the same Choice, Score and Noul questions and returns answers in the same shape. Check cost before switching if your calls carry several questions: Decider bills the state once per question.
- Is Perplexity Decider open source?
- The weights are on Hugging Face under Apache 2.0, and Perplexity says it is fine-tuned from Qwen3.8-27B. Jev has no open weights. If your data cannot leave your own machines, Decider can be self-hosted and Jev cannot.
- Which is more accurate, Jev or Decider?
- In our test they were close. Both scored 86.5% on 96 Banking77 card-intent messages; on 100 AG News articles Decider scored 93% and Jev 91%, a gap inside the noise for a sample this size. Decider was slightly better calibrated: at 0.99 confidence or above it was wrong once, Jev four times.
- Why does Decider cost more on long, multi-question calls?
- Its input tokens scale with the number of questions. On an 8,000-character state we measured 1,392 tokens for one question, 4,192 for three and 6,985 for five, about one full read of the state per question. Jev reads the state once, so its tokens barely move as questions are added.
- What does Decider cost on JevStation?
- 2 credits for a standard evaluation and 6 for a large one, because of the per-question billing. Jev costs 1 and 3. Decider is in the free no-signup trial, so you can compare both on this page without an account.
Related
- Decision modelsJev, Clef, Clef-flash, Perplexity Decider and OpenAI Luna Decisions on the same inputs: accuracy, calibration, cost and latency.
- Decisions API alternativesOpenAI Decisions API without the waitlist, plus Jev, Clef and Perplexity Decider: measured accuracy, cost, and how to port a request.
- Model pricingMonthly cost of five decision models for your volume, text length and question count, from measured token counts.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.