Jev Benchmarks: Every Independent Test So Far
Is Jev accurate? Every independent Jev test to date, from an AI-control pilot to GLiNER and routing tests, with sample sizes, caveats and vendor claims apart.
Jev looks accurate on the few tasks where it has been tested independently, but the evidence is thin: a handful of small pilots, no large human-labelled benchmark, and one clear case where it was badly calibrated. It ranked backdoored code well in an AI-control pilot, beat a zero-shot classifier on two of three text datasets, and routed 40 prompts to the expected tier. Each of those tests is small, and each author says so.
This page collects every independent measurement we could trace, keeps TypeSafe's own numbers in a separate section, and ends with what nobody has measured yet. We run a site on Jev, so we have tried to be stricter with the good results than with the bad ones.
The summary
| Test | Who ran it | Task | Headline result | Sample | Main caveat |
|---|---|---|---|---|---|
| AI-control backdoor monitor | Venkat T, on LessWrong | Rank backdoored code above honest code (one Noul) | AUROC 0.976; about 9 in 10 backdoors caught at a 2% audit budget | 600 backdoors vs 1,200 honest solutions | Toy setting, one dataset, no LLM baseline run; weak raw probabilities |
| jev-benchmarks | AbdelStark, on GitHub (pre-registered) | Zero-shot text classification vs GLiNER | Accuracy 0.91 and 0.87 vs 0.70 and 0.61 on two datasets; tie on the third | 300 examples (100 per dataset) | Badly calibrated on the third dataset; latency not like for like |
| LLM routing | A Classmethod engineer (DevelopersIO) | Pick one of four model tiers | 40 of 40 calls returned the expected tier | 40 calls, one prompt per tier | Author: not an accuracy benchmark |
| Rescoring TypeSafe's examples | jev-on-a-laptop, on GitHub | Agreement with TypeSafe's reference answers | 86.6%, within a frontier cluster of 86–90% | 343 scored question pairs from 20 cases | Reference answers come from frontier models, not people |
| LLM Chess leaderboard | A dev.to author | Play full chess games | Elo about 243, rank about 59; every game completed | 80 games | Not a task Jev is built for |
| agent-assist latency | jev-experiments (built by the Devin agent) | Pick a support macro, nine questions per call | Jev call p50 88 ms, p95 209 ms | 40 messages | Latency only; the LLM baseline was simulated |
| Cost per evaluation | JevStation | Token usage across request sizes | $0.000012 to $0.000377 per evaluation | 11 calls | Cost only; one model build |
The AI-control pilot: strong ranking, weak probabilities
The most thorough test so far is an AI-control pilot on ControlArena's APPS backdoor setting, written up on LessWrong. The whole monitor was one Noul asking whether the code contains a planted backdoor, on about 1,050 tokens of state. The test set was 600 verified backdoors against 600 LLM-written and 600 human-written honest solutions.
- Ranking: AUROC 0.976 against LLM-written honest code and 0.970 against human-written code; 89.5% and 87.8% of backdoors caught at a 2% false-positive rate.
- Consistency: re-scoring the same code five times moved scores by 0.008 on average, and the flagged top 10% overlapped 92.4% between runs.
- Controls: a model that only saw code length scored 0.674, so Jev was not just reading length.
The author is equally clear about the weaknesses. Jev was "quite under-confident", which makes the raw score a good ranking and a poor probability. An attacker who could check the score a few times before submitting got about a third of backdoors through instead of a tenth. Jev caught added branches far more often than edited logic. And the author did not run an LLM baseline, so the comparison with published reasoning-model monitors (0.92 to 0.96) is positioning, not a measurement. The author ran it on about $5 of credit. Jev as a Judge goes through every number.
jev-benchmarks: better than GLiNER on two of three, worse calibrated on one
jev-benchmarks is the only pre-registered comparison. It ran jev-1.13.0 against fastino/gliner2.5-multi-v1 on 100 held-out examples from each of three datasets.
| Dataset | Labels | Jev accuracy | GLiNER accuracy | Jev − GLiNER, 95% CI | Jev coverage at ≤5% error | GLiNER coverage |
|---|---|---|---|---|---|---|
| AG News | 4 | 0.910 | 0.700 | +0.130 to +0.290 | 0.830 | 0.240 |
| Banking77 | 72 | 0.870 | 0.610 | +0.220 to +0.300 | 0.860 | 0.270 |
| DAIR Emotion | 6 | 0.480 | 0.440 | −0.070 to +0.150 | 0.000 | 0.020 |
On emotion labels the accuracy gap is unresolved, and Jev was clearly worse calibrated: Brier 0.846 against 0.668, NLL 5.588 against 1.381, and zero probability on the true label for 16% of examples. The author calls the result deliberately mixed. Latency is not comparable: GLiNER ran on a local CPU and Jev was a hosted API called from France. Jev Alternatives covers what this means when choosing between the two.
The Classmethod routing test: 40 of 40, by the author's own terms a smoke test
A Classmethod engineer tested Jev as the classifier in an LLM router with four tiers: simple, medium, complex and reasoning. All 40 calls, ten per tier, returned the expected tier. Median latency was 0.643 to 0.674 seconds and the cost about $0.000025 to $0.000027 per call. The extreme tiers came back with confidence 1.0; the medium tier sat between 0.57 and 0.67.
The author states the limits: one prompt per tier, Jev called on its own rather than inside the router, and not an exhaustive test of accuracy. The low confidence on the middle tier is the most useful finding, because borderline prompts are where a router saves or loses money. See the LLM router tool for the question set.
jev-on-a-laptop: rescoring TypeSafe's own examples
jev-on-a-laptop is mainly a local reproduction of the idea behind Jev on small open models, with no access to Jev itself. It also rebuilt the public example cases of TypeSafe's four workflow evaluations, 20 cases, and scored each model's answers against TypeSafe's reference: Jev 86.6% (297 of 343 pairs), against 89.2% to 89.8% for Opus, Sol and DeepSeek v4.1 Flash, and 73.8% for a local 7B model. The author reads that as Jev sitting inside the frontier cluster.
Two limits. The reference answers are TypeSafe's, and they come from frontier models rather than people, so this measures agreement, not correctness. And the author notes that an earlier five-case run showed a tie that turned out to be an artifact of the small sample, a useful warning for every test on this page.
Chess and latency: what they do and do not show
A dev.to author ran Jev on the LLM Chess leaderboard: Elo about 243, around rank 59, with every game played to the end and no broken moves. Eighty games cost about $0.12. The score says little about classification; the full games say the typed output kept to the protocol. The same author found Jev weak at generating words one letter at a time.
In jev-experiments, a contact-centre demo built by the Devin coding agent sends nine questions per customer message. Over 40 messages, the Jev call took 88 ms at p50 and 209 ms at p95. That is a latency measurement only, and the "43x faster" comparison in the repository is against a simulated 4-second LLM, which the author labels as simulated.
Our own measurement: cost, not accuracy
We sent 11 real requests across every state size and question count JevStation allows and recorded the tokens Jev reported. The result: about 260 tokens of fixed overhead per request, then roughly 6 characters per token of state and 5 per token of question JSON, for $0.000012 to $0.000377 per evaluation. The method and its limits are in What One Jev Evaluation Costs. It says nothing about accuracy.
Practitioner reports without data
Two reports in TechCrunch are often repeated. A Vercel engineer said replacing an LLM command-safety classifier with Jev gave results five to 18 times faster and with greater accuracy. Bryo AI's CTO found Gemini slightly more accurate at classifying business emails, but 10 to 20 times more expensive. Neither published a dataset or method, so treat them as anecdotes.
One more figure circulates: a "67.8% against 74.1%" result attributed to an independent test by Every. Those numbers match TypeSafe's own published evaluation exactly, and we could not find the independent study, so we do not count it.
Vendor-stated numbers
These come from TypeSafe. They are not independent and should be read with TypeSafe's own caveats.
- Price: $0.042 per million input tokens, output free.
- Latency: 70–500 ms end to end, described as 40–200x faster than frontier LLMs; the homepage figures of 193.6x faster and 444.6x cheaper are, in TypeSafe's words, likely at the higher end of real-world gains.
- Workflow evaluations: Jev averaged 67.8% (61.7% to 76.0% per workflow) at about $0.0004 and 0.4 seconds per case. TypeSafe notes the reference answer is the average of two frontier LLMs, which biases results towards those models, and that its own team built the workflows. Third-party write-ups put a GPT-5.6 model at 67.9% and a stronger one at 74.1% in the same harness, as covered in What a System One Model Is.
- Batching: 13 questions in one call instead of 13 calls, 11.5x cheaper and 9.6x faster.
- Type errors: 0%, which TypeSafe says is not empirical: the schema is guaranteed.
TypeSafe also says it deliberately does not publish results on public benchmarks.
What is still unmeasured
- A large, human-labelled accuracy benchmark. Every accuracy figure here is either a small pilot or agreement with model-made reference answers.
- Calibration on real workloads. One study found it good on two datasets and poor on the third; the other found Jev under-confident. No published calibration curves exist for Jev's own tasks.
- A head-to-head with an LLM on the same inputs, run by someone other than TypeSafe, with the LLM baseline actually executed.
- Non-English accuracy. TypeSafe says other languages are handled less well; nobody has published numbers.
- Stability across model versions. Every test above ran on the only build released so far,
jev-1.13.0. Log themodelfield each response reports, and re-check your thresholds when it changes. - The
confidenceformula, which TypeSafe has not published.
The practical conclusion is the one every author above reaches: measure on your own data. Paste twenty inputs whose answers you know into the playground; the first 200 evaluations are free. For the tasks people are testing, see Jev Use Cases.
Sources
- Backdoor-monitor pilot. Venkat T, LessWrong (independent)
- Jev vs GLiNER. jev-benchmarks (independent, pre-registered)
- Routing test. DevelopersIO (Classmethod) (independent)
- Rescoring TypeSafe's examples. jev-on-a-laptop (independent)
- Chess. dev.to (independent)
- agent-assist latency. jev-experiments (third party)
- Per-evaluation cost. What One Jev Evaluation Costs (our measurement)
- Vercel and Bryo AI reports. TechCrunch via Yahoo (third party)
- Price, latency, speed multiples, eval caveats, no public benchmarks. TypeSafe launch post (vendor)
- Workflow evaluations. evals.typesafe.ai (vendor; figures as quoted in jev-on-a-laptop)
- Batching figures. TypeSafe primitives (vendor)
- Language support, model identity. TypeSafe model docs (vendor)