Jev vs Microsoft-Decision-1: two $0.042 decision models, tested on the same inputs
Microsoft-Decision-1 launched on 2026-10-09 and is already on OpenRouter. It takes the same request as Jev, a state plus typed questions, and returns probabilities, at the same list price: $0.042 per million input tokens. In our run on 2026-10-10 it matched Jev on accuracy (87.4% against 86.5% on Banking77 card intents, 91% against 91% on AG News), was better calibrated, and used fewer tokens on every request we measured. The catch is rate limits: Azure throttled it once during a four-way parallel run.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Route a ticket to the right team, rate its urgency and flag churn risk.
Edit freely — the model answers the scenario's questions about whatever you put here.
What Jev actually returned
This is the unedited output of the example above, captured from the live API.
Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing with 'invalid redirect'. I'm losing sales every hour this is down. Please help ASAP.
category
choiceWhich team should handle this support ticket?
integrations
- integrations99%
- billing1%
- bug0%
- account0%
Confidence 99%
urgency
scoreHow urgent is this ticket?
- 4 · Critical73%
- 3 · Urgent27%
- 0 · Can wait0%
- 1 · Normal0%
- 2 · Soon0%
Confidence 77%
churn risk
noulIs the customer at risk of churning?
Captured 2026-10-09 · 0.89 s · jev-1.13.0
At a glance
(vendor) is the company selling the model; (ours) is a call we made through OpenRouter (Jev on 2026-10-07, Decision-1 on 2026-10-10).
| Jev | Microsoft-Decision-1 | |
|---|---|---|
| Made by | TypeSafe | Microsoft |
| Model id | jev-latest | microsoft/microsoft-decision-1 |
| Weights | Closed | Closed, Foundry and OpenRouter only |
| Base model | Not disclosed | Qwen3.5-9B (vendor) |
| List price, input | $0.042 / M tokens | $0.042 / M tokens |
| Output tokens | Free | Free |
| Context | ~32k tokens, shared with the longest question (vendor) | 32,768 tokens (vendor) |
| Images | No | No |
| Input tokens, our fit | ≈ 260 + 0.163 × characters + 43 × questions | ≈ 0.163 × characters + 35 × questions |
| Median latency | 412 ms (ours) | 473–491 ms (ours) |
| Banking77 card intents | 86.5% (ours) | 87.4% (ours) |
| AG News topics | 91% (ours) | 91% (ours) |
| Credits on JevStation | 1 / 3 | 1 / 3 |
Credits are standard / large; an evaluation is large above 8,000 characters of state or five questions.
Accuracy and calibration
We sent both models the same labelled inputs, zero-shot, with one line of description per option. The method is on Decision models compared, which runs the other models on the same data.
| Task | Jev | Decision-1 |
|---|---|---|
| Banking77, 12 card intents (96) | 86.5% | 87.4% (95 answered) |
| AG News, 4 topics (100) | 91% | 91% |
| Banking77: answers at ≥ 0.9 confidence | 81% of inputs, 94.9% right | 67% of inputs, 100% right |
| Wrong at ≥ 0.99 confidence | 4 of 196 | 2 of 195 |
| Expected calibration error, AG News | 0.086 | 0.053 |
Accuracy is a tie. Calibration is where Decision-1 looks better: when it was very sure it was almost always right, and it did not sit at 1.00 as often (74 of 195 answers at 0.99 or above, against 136 of 196 for Jev), which leaves room to separate easy cases from risky ones. The trade is coverage: fewer answers clear a 0.9 threshold, so more go to a human.
Decision-1 also found a single cancel request hidden anywhere in 20,000 characters of neutral text (0.998 at the very end).
Tokens, latency and limits
The list price is identical, so the bill follows the token count. Decision-1 counts fewer tokens for the same request, and like Jev it reads the state once. We fitted it on nine requests (state of 100, 2,000 and 8,000 characters, with one, three and five questions):
| Request | Jev tokens | Decision-1 tokens | Cheaper at list price |
|---|---|---|---|
| 1 question, short ticket (~300 chars) | ~350 | ~85 | Decision-1, about 4× |
| 3 questions, 1,500 characters | ~640 | ~350 | Decision-1, about 2× |
| 5 questions, 8,000 characters | ~1,780 | ~1,470 | Decision-1, about 1.2× |
| 12 long questions, 24,000 characters | 8,981 | 8,150 | Decision-1, about 1.1× |
Latency was about the same as Jev, with a heavier tail: p50 near 480 ms, p90 up to 1.4 s on Banking77. Azure’s rate limit is lower than we expected: with four requests in flight, one of 96 came back as HTTP 429. Retry with backoff if you run bursts; JevStation does this for you.
When to pick which
- Short, single-question calls at volume: Decision-1, the fewest tokens we measured (Perplexity Decider is close, at about $0.007 per thousand AG News calls against $0.005).
- A confidence you can act on automatically: Decision-1 was better calibrated in our run; confirm on your own data.
- Maximum coverage at a fixed threshold: Jev, which clears 0.9 more often.
- Self-hosting or data residency: neither. Use Perplexity Decider, which is open-weight.
- Images or inputs over ~30k tokens: Perplexity Decider.
- A vendor-backed SLA on Azure: Decision-1.
Try both
Run Decision-1 above without an account, or sign in and use the side-by-side runner to send the same text and questions to Jev and Decision-1 at once. The pricing calculator shows the cost of each for your own numbers.
Sources
- Decision-1 model, price and context. OpenRouter model page (vendor)
- Launch, base model and benchmark claims. MarkTechPost and OrcaRouter (independent)
- Jev price and context. TypeSafe model docs (vendor)
- Accuracy, calibration, latency and token counts. Our runs through OpenRouter on 2026-10-07 (Jev) and 2026-10-10 (Decision-1) (ours)
JevStation is independent. Jev and Microsoft-Decision-1 are names of their respective owners, TypeSafe and Microsoft; we are not affiliated with either.
Frequently asked questions
- Is Microsoft-Decision-1 a drop-in replacement for Jev?
- For the request format, yes. Through OpenRouter’s System One endpoint or JevStation it accepts the same state and the same Choice, Score and Noul questions and returns answers in the same shape. Two differences: it takes text only, and its context is about 32k tokens.
- Can I download or self-host Microsoft-Decision-1?
- No. Microsoft distributes it only through Microsoft Foundry (and OpenRouter, which routes to Azure). There are no published weights, although Microsoft says it is post-trained from Qwen3.5-9B. If you need self-hosting, Perplexity Decider is the open-weight option.
- Which is more accurate, Jev or Decision-1?
- In our test they were level. Decision-1 scored 87.4% on 95 answered Banking77 card-intent messages (one request was rate limited) against Jev’s 86.5% on 96; on 100 AG News articles both scored 91%. That is inside the noise for a sample this size. Microsoft reports the highest accuracy across 36 benchmarks, but those figures are its own and have not been independently reproduced.
- Is Decision-1 better calibrated than Jev?
- In our run, yes. On Banking77 every answer at 0.9 confidence or above was right (64 of 64), against 94.9% for Jev. It was wrong twice at 0.99 or above out of 74 such answers; Jev was wrong four times, and it reported 1.00 on 136 of 196 inputs. Treat it as a sign, not a guarantee, and check on your own data.
- What does Decision-1 cost on JevStation?
- 1 credit for a standard evaluation and 3 for a large one, the same as Jev. It is in the free no-signup trial, so you can compare both on this page without an account.
Related
- Decision modelsJev, Clef, Clef-flash, Perplexity Decider and OpenAI Luna Decisions on the same inputs: accuracy, calibration, cost and latency.
- Jev vs DeciderPerplexity Decider against Jev: same contract, measured accuracy, calibration, latency, and per-question billing.
- Model pricingMonthly cost of five decision models for your volume, text length and question count, from measured token counts.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.