Back to blog

Jev Benchmarks: Every Independent Test So Far

Is Jev accurate? Every independent Jev test to date, from an AI-control pilot to GLiNER and routing tests, with sample sizes, caveats and vendor claims apart.

JevStationJevStation

Jev looks accurate on the few tasks where it has been tested independently, but the evidence is thin: a handful of small pilots, no large human-labelled benchmark, and one clear case where it was badly calibrated. It ranked backdoored code well in an AI-control pilot, beat a zero-shot classifier on two of three text datasets, and routed 40 prompts to the expected tier. Each of those tests is small, and each author says so.

This page collects every independent measurement we could trace, keeps TypeSafe's own numbers in a separate section, and ends with what nobody has measured yet. We run a site on Jev, so we have tried to be stricter with the good results than with the bad ones.

The summary

TestWho ran itTaskHeadline resultSampleMain caveat
AI-control backdoor monitorVenkat T, on LessWrongRank backdoored code above honest code (one Noul)AUROC 0.976; about 9 in 10 backdoors caught at a 2% audit budget600 backdoors vs 1,200 honest solutionsToy setting, one dataset, no LLM baseline run; weak raw probabilities
jev-benchmarksAbdelStark, on GitHub (pre-registered)Zero-shot text classification vs GLiNERAccuracy 0.91 and 0.87 vs 0.70 and 0.61 on two datasets; tie on the third300 examples (100 per dataset)Badly calibrated on the third dataset; latency not like for like
LLM routingA Classmethod engineer (DevelopersIO)Pick one of four model tiers40 of 40 calls returned the expected tier40 calls, one prompt per tierAuthor: not an accuracy benchmark
Rescoring TypeSafe's examplesjev-on-a-laptop, on GitHubAgreement with TypeSafe's reference answers86.6%, within a frontier cluster of 86–90%343 scored question pairs from 20 casesReference answers come from frontier models, not people
LLM Chess leaderboardA dev.to authorPlay full chess gamesElo about 243, rank about 59; every game completed80 gamesNot a task Jev is built for
agent-assist latencyjev-experiments (built by the Devin agent)Pick a support macro, nine questions per callJev call p50 88 ms, p95 209 ms40 messagesLatency only; the LLM baseline was simulated
Cost per evaluationJevStationToken usage across request sizes$0.000012 to $0.000377 per evaluation11 callsCost only; one model build

The AI-control pilot: strong ranking, weak probabilities

The most thorough test so far is an AI-control pilot on ControlArena's APPS backdoor setting, written up on LessWrong. The whole monitor was one Noul asking whether the code contains a planted backdoor, on about 1,050 tokens of state. The test set was 600 verified backdoors against 600 LLM-written and 600 human-written honest solutions.

  • Ranking: AUROC 0.976 against LLM-written honest code and 0.970 against human-written code; 89.5% and 87.8% of backdoors caught at a 2% false-positive rate.
  • Consistency: re-scoring the same code five times moved scores by 0.008 on average, and the flagged top 10% overlapped 92.4% between runs.
  • Controls: a model that only saw code length scored 0.674, so Jev was not just reading length.

The author is equally clear about the weaknesses. Jev was "quite under-confident", which makes the raw score a good ranking and a poor probability. An attacker who could check the score a few times before submitting got about a third of backdoors through instead of a tenth. Jev caught added branches far more often than edited logic. And the author did not run an LLM baseline, so the comparison with published reasoning-model monitors (0.92 to 0.96) is positioning, not a measurement. The author ran it on about $5 of credit. Jev as a Judge goes through every number.

jev-benchmarks: better than GLiNER on two of three, worse calibrated on one

jev-benchmarks is the only pre-registered comparison. It ran jev-1.13.0 against fastino/gliner2.5-multi-v1 on 100 held-out examples from each of three datasets.

DatasetLabelsJev accuracyGLiNER accuracyJev − GLiNER, 95% CIJev coverage at ≤5% errorGLiNER coverage
AG News40.9100.700+0.130 to +0.2900.8300.240
Banking77720.8700.610+0.220 to +0.3000.8600.270
DAIR Emotion60.4800.440−0.070 to +0.1500.0000.020

On emotion labels the accuracy gap is unresolved, and Jev was clearly worse calibrated: Brier 0.846 against 0.668, NLL 5.588 against 1.381, and zero probability on the true label for 16% of examples. The author calls the result deliberately mixed. Latency is not comparable: GLiNER ran on a local CPU and Jev was a hosted API called from France. Jev Alternatives covers what this means when choosing between the two.

The Classmethod routing test: 40 of 40, by the author's own terms a smoke test

A Classmethod engineer tested Jev as the classifier in an LLM router with four tiers: simple, medium, complex and reasoning. All 40 calls, ten per tier, returned the expected tier. Median latency was 0.643 to 0.674 seconds and the cost about $0.000025 to $0.000027 per call. The extreme tiers came back with confidence 1.0; the medium tier sat between 0.57 and 0.67.

The author states the limits: one prompt per tier, Jev called on its own rather than inside the router, and not an exhaustive test of accuracy. The low confidence on the middle tier is the most useful finding, because borderline prompts are where a router saves or loses money. See the LLM router tool for the question set.

jev-on-a-laptop: rescoring TypeSafe's own examples

jev-on-a-laptop is mainly a local reproduction of the idea behind Jev on small open models, with no access to Jev itself. It also rebuilt the public example cases of TypeSafe's four workflow evaluations, 20 cases, and scored each model's answers against TypeSafe's reference: Jev 86.6% (297 of 343 pairs), against 89.2% to 89.8% for Opus, Sol and DeepSeek v4.1 Flash, and 73.8% for a local 7B model. The author reads that as Jev sitting inside the frontier cluster.

Two limits. The reference answers are TypeSafe's, and they come from frontier models rather than people, so this measures agreement, not correctness. And the author notes that an earlier five-case run showed a tie that turned out to be an artifact of the small sample, a useful warning for every test on this page.

Chess and latency: what they do and do not show

A dev.to author ran Jev on the LLM Chess leaderboard: Elo about 243, around rank 59, with every game played to the end and no broken moves. Eighty games cost about $0.12. The score says little about classification; the full games say the typed output kept to the protocol. The same author found Jev weak at generating words one letter at a time.

In jev-experiments, a contact-centre demo built by the Devin coding agent sends nine questions per customer message. Over 40 messages, the Jev call took 88 ms at p50 and 209 ms at p95. That is a latency measurement only, and the "43x faster" comparison in the repository is against a simulated 4-second LLM, which the author labels as simulated.

Our own measurement: cost, not accuracy

We sent 11 real requests across every state size and question count JevStation allows and recorded the tokens Jev reported. The result: about 260 tokens of fixed overhead per request, then roughly 6 characters per token of state and 5 per token of question JSON, for $0.000012 to $0.000377 per evaluation. The method and its limits are in What One Jev Evaluation Costs. It says nothing about accuracy.

Practitioner reports without data

Two reports in TechCrunch are often repeated. A Vercel engineer said replacing an LLM command-safety classifier with Jev gave results five to 18 times faster and with greater accuracy. Bryo AI's CTO found Gemini slightly more accurate at classifying business emails, but 10 to 20 times more expensive. Neither published a dataset or method, so treat them as anecdotes.

One more figure circulates: a "67.8% against 74.1%" result attributed to an independent test by Every. Those numbers match TypeSafe's own published evaluation exactly, and we could not find the independent study, so we do not count it.

Vendor-stated numbers

These come from TypeSafe. They are not independent and should be read with TypeSafe's own caveats.

  • Price: $0.042 per million input tokens, output free.
  • Latency: 70–500 ms end to end, described as 40–200x faster than frontier LLMs; the homepage figures of 193.6x faster and 444.6x cheaper are, in TypeSafe's words, likely at the higher end of real-world gains.
  • Workflow evaluations: Jev averaged 67.8% (61.7% to 76.0% per workflow) at about $0.0004 and 0.4 seconds per case. TypeSafe notes the reference answer is the average of two frontier LLMs, which biases results towards those models, and that its own team built the workflows. Third-party write-ups put a GPT-5.6 model at 67.9% and a stronger one at 74.1% in the same harness, as covered in What a System One Model Is.
  • Batching: 13 questions in one call instead of 13 calls, 11.5x cheaper and 9.6x faster.
  • Type errors: 0%, which TypeSafe says is not empirical: the schema is guaranteed.

TypeSafe also says it deliberately does not publish results on public benchmarks.

What is still unmeasured

  • A large, human-labelled accuracy benchmark. Every accuracy figure here is either a small pilot or agreement with model-made reference answers.
  • Calibration on real workloads. One study found it good on two datasets and poor on the third; the other found Jev under-confident. No published calibration curves exist for Jev's own tasks.
  • A head-to-head with an LLM on the same inputs, run by someone other than TypeSafe, with the LLM baseline actually executed.
  • Non-English accuracy. TypeSafe says other languages are handled less well; nobody has published numbers.
  • Stability across model versions. Every test above ran on the only build released so far, jev-1.13.0. Log the model field each response reports, and re-check your thresholds when it changes.
  • The confidence formula, which TypeSafe has not published.

The practical conclusion is the one every author above reaches: measure on your own data. Paste twenty inputs whose answers you know into the playground; the first 200 evaluations are free. For the tasks people are testing, see Jev Use Cases.

Sources