Jev vs Perplexity Decider: the two cheapest decision models, tested on the same inputs

Perplexity Decider (pplx-decider-v1-27b) and Jev take the same request, a state plus typed questions, and return probabilities. They are also the two cheapest per token: $0.04 and $0.042 per million input tokens. In our run on 2026-10-07 they tied on Banking77 card intents (86.5%) and Decider edged ahead on AG News (93% against 91%). The difference that matters is billing: Decider reads the state again for every question, so it is the cheaper model for one question and the more expensive one for several on long text.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Route a ticket to the right team, rate its urgency and flag churn risk.

166 / 2,000

Edit freely — the model answers the scenario's questions about whatever you put here.

Model

At a glance

(vendor) is the company selling the model; (ours) is a call we made on 2026-10-07 through OpenRouter.

JevPerplexity Decider
Made byTypeSafePerplexity
Model idjev-latestpplx-decider-v1-27b
WeightsClosedOpen, Apache 2.0
Base modelNot disclosedQwen3.8-27B (vendor)
List price, input$0.042 / M tokens$0.04 / M tokens
Output tokensFreeFree
Context~32k tokens, shared with the longest question (vendor)262k tokens (vendor)
ImagesNoYes
Input tokens, our fit≈ 260 + 0.163 × characters + 43 × questions≈ questions × (84 + 0.163 × characters)
Median latency412 ms (ours)493 ms (ours)
Banking77 card intents86.5% (ours)86.5% (ours)
AG News topics91% (ours)93% (ours)
Credits on JevStation1 / 32 / 6

Credits are standard / large; an evaluation is large above 8,000 characters of state or five questions.

Accuracy and calibration

We sent both models the same 196 labelled inputs, zero-shot, with one line of description per option. The method is on Decision models compared, which runs all five models on the same data.

TaskJevDecider
Banking77, 12 card intents (96)86.5%86.5%
AG News, 4 topics (100)91%93%
Banking77: answers at ≥ 0.9 confidence81% of inputs, 94.9% right77% of inputs, 98.6% right
Wrong at ≥ 0.99 confidence (of 196)41
Expected calibration error, AG News0.0860.052

The accuracy is a tie at this sample size. Calibration is where Decider pulled ahead: when it was very sure, it was nearly always right, so a threshold set on its confidence lets fewer mistakes through. Jev gave a confidence of 1.00 (to two decimals) on 136 of the 196 inputs, which leaves little room to separate easy cases from risky ones.

Both models found a single cancel request hidden at the very end of 20,000 characters of neutral text, so neither truncated the state in our test.

The cost difference is in the token count, not the price

The list prices are almost the same. The bills are not, because the two models count tokens differently. We fitted both on nine requests each (state of 100, 2,000 and 8,000 characters, with one, three and five questions):

RequestJev tokensDecider tokensCheaper at list price
1 question, short ticket (~300 chars)~350~130Decider
3 questions, 1,500 characters~640~990Jev
5 questions, 8,000 characters~1,7806,985Jev, by about 4×
12 long questions, 24,000 characters8,98151,882Jev, by about 6×

Decider’s count grows with questions × state length, as if each question were a separate call. Jev’s grows with state length plus a small amount per question. If your workload is one question per call, such as a router or a single guardrail, Decider is the cheapest model we have measured. If you ask several things about the same document, Jev is cheaper, even though its list price is higher.

On JevStation that is why Decider costs 2 credits rather than 1. The pricing calculator shows both for your own numbers.

When to pick which

  • One question per call at high volume: Decider.
  • Several questions on the same text: Jev.
  • You need images judged: Decider. Jev does not take images.
  • You need a confidence you can act on automatically: Decider was better calibrated in our run; confirm on your own data.
  • Self-hosting or data residency: Decider, open-weight. Jev is API-only.
  • Inputs longer than ~30k tokens: Decider, with a 262k context.

Try both

Run Decider above without an account, or sign in and use the side-by-side runner to send the same text and questions to Jev and Decider at once.

Sources

JevStation is independent. Jev and Decider are names of their respective owners, TypeSafe and Perplexity; we are not affiliated with either.

Frequently asked questions

Is Perplexity Decider a drop-in replacement for Jev?
For the request format, yes. Through OpenRouter’s System One endpoint or JevStation, Decider accepts the same state and the same Choice, Score and Noul questions and returns answers in the same shape. Check cost before switching if your calls carry several questions: Decider bills the state once per question.
Is Perplexity Decider open source?
The weights are on Hugging Face under Apache 2.0, and Perplexity says it is fine-tuned from Qwen3.8-27B. Jev has no open weights. If your data cannot leave your own machines, Decider can be self-hosted and Jev cannot.
Which is more accurate, Jev or Decider?
In our test they were close. Both scored 86.5% on 96 Banking77 card-intent messages; on 100 AG News articles Decider scored 93% and Jev 91%, a gap inside the noise for a sample this size. Decider was slightly better calibrated: at 0.99 confidence or above it was wrong once, Jev four times.
Why does Decider cost more on long, multi-question calls?
Its input tokens scale with the number of questions. On an 8,000-character state we measured 1,392 tokens for one question, 4,192 for three and 6,985 for five, about one full read of the state per question. Jev reads the state once, so its tokens barely move as questions are added.
What does Decider cost on JevStation?
2 credits for a standard evaluation and 6 for a large one, because of the per-question billing. Jev costs 1 and 3. Decider is in the free no-signup trial, so you can compare both on this page without an account.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.