Jev vs an LLM classifier: when a typed decision beats a prompt

For a decision whose answers you can list in advance and that you make over and over, Jev is the better classifier on cost and speed, and roughly level on accuracy in the tests that exist. A general LLM is still the better tool when you need the reasoning written out, arithmetic or dates, an open-ended answer, or the last few points of accuracy.

At a glance

The comparison here is between Jev and a general LLM prompted to pick a label, with or without structured output. For zero-shot classifiers and models you train yourself, see Jev vs open models and the overview in Jev Alternatives.

JevPrompted LLM (GPT / Claude class)
OutputA typed value from the options you declaredText, parsed into a label
ProbabilitiesReturned for every answerOnly if you ask, as generated text
Price (vendor-stated)$0.042 / M input tokens, output free$0.20–$10 / M input tokens, output about 5x input
Cost per sample, TypeSafe evals~$0.0004$0.0304–$0.0836
Latency, TypeSafe evals0.4 s10.1–23.3 s
Accuracy, TypeSafe evals67.8%67.9–74.1%
Written reasoningNoneYes
Counting, arithmetic, datesDocumented weaknessBetter suited
Open-ended answersNoYes

The evaluation rows come from TypeSafe’s four published workflows as reported by a third-party write-up. They are vendor numbers, and the reference answers in them are the average of two frontier LLMs, not human labels. Prices checked 2026-09-19.

Cost

Jev bills input tokens only, and a classification call is small. We measured real evaluations at $0.000012 to $0.000377, with a typical three-question call at about $0.00002 (the measurement). The LLM side depends on the model, and output tokens are billed there too.

Two independent reports point the same way. The CTO of Bryo AI told TechCrunch that in his test of business-email classification Gemini was slightly more accurate but 10 to 20 times more expensive. A DevelopersIO router test measured $0.000025–$0.000027 per Jev call.

Batching widens the gap. Jev answers every question in a request in parallel, and TypeSafe’s docs say 13 questions in one call is 11.5x cheaper and 9.6x faster than 13 separate calls. With an LLM, asking more questions usually means more output tokens or more calls.

Latency

TypeSafe claims 70–500 ms end to end, against a 3–329 second range for LLMs. That is vendor-stated. Independent numbers are smaller and more useful:

  • Model routing: DevelopersIO measured a median of 0.643–0.674 s per Jev call, about 3x faster than a Gemini 3.5 Flash classifier and 10–11x faster than a DeepSeek V4 Flash one. The author notes it called TypeSafe standalone, with one sample per tier.
  • Code monitoring: a LessWrong AI-control pilot recorded a median latency of 0.34 s.
  • Command safety review: TechCrunch reports that Vercel got results 5 to 18 times faster, and more accurate, after replacing an OpenAI Luna 5.6 classifier with Jev.

Output shape: typed values vs parsed text

An LLM with structured output still generates text; the schema constrains the shape, and your code parses and validates it. Jev does not generate text. You declare the options, and the answer is one of them, with a probability for each. It cannot return a label you did not list.

TypeSafe says Jev makes no type errors and is candid that this is guaranteed by schema matching, not measured. The consequence shows up in unrelated tests: on the LLM Chess leaderboard, Jev completed every game with perfect protocol compliance. The caveat is equally plain — a typed answer can still be the wrong answer.

Accuracy: what has been measured

No independent, human-labelled head-to-head between Jev and an LLM classifier exists in our sources. What does exist:

  • TypeSafe’s workflow evals (vendor): Jev 67.8%, GPT-5.6 Terra 67.9%, GPT-5.6 Sol 74.1%. TypeSafe itself warns these were built by its own team, that the reference answers favour OpenAI and Anthropic models, and that the LLMs ran through its wrapper.
  • An independent re-scoring by jev-on-a-laptop of 343 question pairs from TypeSafe’s public examples: Opus 89.8%, Sol 89.2%, Jev 86.6%. The author’s reading is that Jev sits in the frontier cluster at a fraction of the price.
  • Anecdotes: Bryo AI found Gemini slightly more accurate; Vercel found Jev more accurate than the classifier it replaced.

Read together: parity, give or take a few points, depending on the task. Test it on your own labelled data before relying on either.

Calibration and consistency

This is the part an LLM cannot easily match. Jev is trained to return probabilities that mean what they say across groups of predictions, and TypeSafe describes LLMs as overconfident and inconsistent (a vendor claim). The independent evidence is mixed, and worth knowing:

  • Consistency is high. In the LessWrong pilot, re-scoring the same code moved scores by 0.008 on average, and 92% of the flagged set stayed the same across calls.
  • Calibration varies by task. A pre-registered study found Jev well calibrated on two datasets and badly on a third (details in Jev vs open models).
  • Jev can be under-confident. The LessWrong author called the raw score a great ranking and a poor probability.
  • Equivalent questions disagree. TypeSafe documents that a Noul and the equivalent yes/no Choice can give different probabilities, and that a question and its negation can sum to more than 1. Tune thresholds on the exact questions you ship.

When an LLM is the better choice

  • You need the reasoning written down, for a reviewer, an audit trail or a regulated decision.
  • The decision involves counting, arithmetic or comparing dates. TypeSafe’s jaggedness notes say Jev is not a calculator.
  • You cannot list the possible answers in advance, or you need text back.
  • Your inputs are mostly not English; TypeSafe says other languages, including CJK scripts, are handled but not as well.
  • A few points of accuracy matter more than cost and latency.

When Jev is the better choice

  • The answer space is known and the same decision repeats at volume.
  • You want to act on confidence: automatically above one threshold, by escalation below it.
  • Latency matters, as in a pre-execution check or a live support queue.
  • You ask several questions about the same input and want them answered in one call.

Using both

The pattern most of the evidence points to is a split, not a swap. Let the LLM extract or draft, and let Jev verify with a typed yes/no and a probability. Or put Jev first: act on the answers it is confident about, and send the middle band to an LLM or a person. Jev as a judge and content moderation show both shapes, and you can try your own questions in the playground.

よくある質問

Is Jev more accurate than GPT for classification?
Not on the evidence so far. In TypeSafe’s own workflow evaluations Jev scored 67.8% against 67.9% for GPT-5.6 Terra and 74.1% for GPT-5.6 Sol. An independent re-scoring of TypeSafe’s public examples put Jev at 86.6% against 89.2–89.8% for frontier models. The case for Jev is similar accuracy at a fraction of the cost and time, not higher accuracy.
How much cheaper is Jev than an LLM classifier?
In TypeSafe’s evaluations Jev cost about $0.0004 per sample against $0.0304 for GPT-5.6 Terra and $0.0836 for GPT-5.6 Sol (vendor-stated). We measured a single Jev evaluation at $0.000012 to $0.000377 in tokens.
Can’t I just ask the LLM for a confidence score?
You can, but the number is more text the model generated. Jev’s probabilities come from the distribution over the options you declared, which is why they can be thresholded. They are not perfect: independent tests found Jev well calibrated on some tasks, badly on one, and under-confident as a monitor.
Should I replace my LLM with Jev?
Replace the classification calls, not the model. Jev does not generate text, so the LLM stays for drafting, reasoning and extraction. A common split is to let Jev decide what it is confident about and send the rest to the LLM or a person.

関連

パイプラインに組み込む

新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。