Jev vs Microsoft-Decision-1: two $0.042 decision models, tested on the same inputs

Microsoft-Decision-1 launched on 2026-10-09 and is already on OpenRouter. It takes the same request as Jev, a state plus typed questions, and returns probabilities, at the same list price: $0.042 per million input tokens. In our run on 2026-10-10 it matched Jev on accuracy (87.4% against 86.5% on Banking77 card intents, 91% against 91% on AG News), was better calibrated, and used fewer tokens on every request we measured. The catch is rate limits: Azure throttled it once during a four-way parallel run.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Route a ticket to the right team, rate its urgency and flag churn risk.

166 / 2,000

Edit freely — the model answers the scenario's questions about whatever you put here.

Model

What Jev actually returned

This is the unedited output of the example above, captured from the live API.

Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing with 'invalid redirect'. I'm losing sales every hour this is down. Please help ASAP.

category

choice

Which team should handle this support ticket?

integrations

  • integrations99%
  • billing1%
  • bug0%
  • account0%

Confidence 99%

urgency

score

How urgent is this ticket?

3.73/4Critical
  • 4 · Critical73%
  • 3 · Urgent27%
  • 0 · Can wait0%
  • 1 · Normal0%
  • 2 · Soon0%

Confidence 77%

churn risk

noul

Is the customer at risk of churning?

Yes80%
Yes 80%No 20%

Captured 2026-10-09 · 0.89 s · jev-1.13.0

At a glance

(vendor) is the company selling the model; (ours) is a call we made through OpenRouter (Jev on 2026-10-07, Decision-1 on 2026-10-10).

JevMicrosoft-Decision-1
Made byTypeSafeMicrosoft
Model idjev-latestmicrosoft/microsoft-decision-1
WeightsClosedClosed, Foundry and OpenRouter only
Base modelNot disclosedQwen3.5-9B (vendor)
List price, input$0.042 / M tokens$0.042 / M tokens
Output tokensFreeFree
Context~32k tokens, shared with the longest question (vendor)32,768 tokens (vendor)
ImagesNoNo
Input tokens, our fit≈ 260 + 0.163 × characters + 43 × questions≈ 0.163 × characters + 35 × questions
Median latency412 ms (ours)473–491 ms (ours)
Banking77 card intents86.5% (ours)87.4% (ours)
AG News topics91% (ours)91% (ours)
Credits on JevStation1 / 31 / 3

Credits are standard / large; an evaluation is large above 8,000 characters of state or five questions.

Accuracy and calibration

We sent both models the same labelled inputs, zero-shot, with one line of description per option. The method is on Decision models compared, which runs the other models on the same data.

TaskJevDecision-1
Banking77, 12 card intents (96)86.5%87.4% (95 answered)
AG News, 4 topics (100)91%91%
Banking77: answers at ≥ 0.9 confidence81% of inputs, 94.9% right67% of inputs, 100% right
Wrong at ≥ 0.99 confidence4 of 1962 of 195
Expected calibration error, AG News0.0860.053

Accuracy is a tie. Calibration is where Decision-1 looks better: when it was very sure it was almost always right, and it did not sit at 1.00 as often (74 of 195 answers at 0.99 or above, against 136 of 196 for Jev), which leaves room to separate easy cases from risky ones. The trade is coverage: fewer answers clear a 0.9 threshold, so more go to a human.

Decision-1 also found a single cancel request hidden anywhere in 20,000 characters of neutral text (0.998 at the very end).

Tokens, latency and limits

The list price is identical, so the bill follows the token count. Decision-1 counts fewer tokens for the same request, and like Jev it reads the state once. We fitted it on nine requests (state of 100, 2,000 and 8,000 characters, with one, three and five questions):

RequestJev tokensDecision-1 tokensCheaper at list price
1 question, short ticket (~300 chars)~350~85Decision-1, about 4×
3 questions, 1,500 characters~640~350Decision-1, about 2×
5 questions, 8,000 characters~1,780~1,470Decision-1, about 1.2×
12 long questions, 24,000 characters8,9818,150Decision-1, about 1.1×

Latency was about the same as Jev, with a heavier tail: p50 near 480 ms, p90 up to 1.4 s on Banking77. Azure’s rate limit is lower than we expected: with four requests in flight, one of 96 came back as HTTP 429. Retry with backoff if you run bursts; JevStation does this for you.

When to pick which

  • Short, single-question calls at volume: Decision-1, the fewest tokens we measured (Perplexity Decider is close, at about $0.007 per thousand AG News calls against $0.005).
  • A confidence you can act on automatically: Decision-1 was better calibrated in our run; confirm on your own data.
  • Maximum coverage at a fixed threshold: Jev, which clears 0.9 more often.
  • Self-hosting or data residency: neither. Use Perplexity Decider, which is open-weight.
  • Images or inputs over ~30k tokens: Perplexity Decider.
  • A vendor-backed SLA on Azure: Decision-1.

Try both

Run Decision-1 above without an account, or sign in and use the side-by-side runner to send the same text and questions to Jev and Decision-1 at once. The pricing calculator shows the cost of each for your own numbers.

Sources

  • Decision-1 model, price and context. OpenRouter model page (vendor)
  • Launch, base model and benchmark claims. MarkTechPost and OrcaRouter (independent)
  • Jev price and context. TypeSafe model docs (vendor)
  • Accuracy, calibration, latency and token counts. Our runs through OpenRouter on 2026-10-07 (Jev) and 2026-10-10 (Decision-1) (ours)

JevStation is independent. Jev and Microsoft-Decision-1 are names of their respective owners, TypeSafe and Microsoft; we are not affiliated with either.

Frequently asked questions

Is Microsoft-Decision-1 a drop-in replacement for Jev?
For the request format, yes. Through OpenRouter’s System One endpoint or JevStation it accepts the same state and the same Choice, Score and Noul questions and returns answers in the same shape. Two differences: it takes text only, and its context is about 32k tokens.
Can I download or self-host Microsoft-Decision-1?
No. Microsoft distributes it only through Microsoft Foundry (and OpenRouter, which routes to Azure). There are no published weights, although Microsoft says it is post-trained from Qwen3.5-9B. If you need self-hosting, Perplexity Decider is the open-weight option.
Which is more accurate, Jev or Decision-1?
In our test they were level. Decision-1 scored 87.4% on 95 answered Banking77 card-intent messages (one request was rate limited) against Jev’s 86.5% on 96; on 100 AG News articles both scored 91%. That is inside the noise for a sample this size. Microsoft reports the highest accuracy across 36 benchmarks, but those figures are its own and have not been independently reproduced.
Is Decision-1 better calibrated than Jev?
In our run, yes. On Banking77 every answer at 0.9 confidence or above was right (64 of 64), against 94.9% for Jev. It was wrong twice at 0.99 or above out of 74 such answers; Jev was wrong four times, and it reported 1.00 on 136 of 196 inputs. Treat it as a sign, not a guarantee, and check on your own data.
What does Decision-1 cost on JevStation?
1 credit for a standard evaluation and 3 for a large one, the same as Jev. It is in the free no-signup trial, so you can compare both on this page without an account.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.