Jev vs open models: what you can run yourself, and what it costs you
No open-weight Jev appears in any source we track. The open projects that exist copy its interface — typed questions, constrained answers — onto small local models, and the one that measured itself scored about 13 points below Jev, with confidence that did not flag its own errors. Run your own model when the data cannot leave your machines or you have a large, stable labelled set; otherwise hosted Jev is the stronger default.
At a glance
"Open-source Jev" gets searched as if it were one thing. In our sources it is four different things, only one of which has been measured against Jev by someone independent. For the short overview of every alternative, including general LLMs, start with Jev Alternatives; this page goes deeper on the options you can run yourself.
| Option | What it is | Runs where | Measured against Jev? |
|---|---|---|---|
| Jev | TypeSafe’s hosted System One model | TypeSafe’s API, or resellers | — |
| jev-on-a-laptop | Jev-style constrained decoding on stock 1.5B–8B models | Your laptop | Yes, on TypeSafe’s public examples |
| open-jev | Typed JSON inference with DiffusionGemma | Your hardware | Not in our sources |
| GLiNER 2.5 | A zero-shot classifier | Your hardware, CPU is enough | Yes, pre-registered pilot |
| Your own classifier | A model you fine-tune on your labels | Your hardware | No public benchmark |
Why there is nothing to download
What makes Jev Jev is the training, and none of it is public. TypeSafe trains it with RLCD (Reinforcement Learning for Calibrated Decisions) on synthetic data it makes in-house, and the same weights serve every account — there is no per-customer fine-tuning or LoRA. Public material does not give the parameter count, the architecture, the full RLCD recipe or calibration curves (what a System One model is).
What can be copied is the interface: declare the options, force the model to answer with one of them, read the probability off the distribution. That is what both open projects below do. It gets you the shape of Jev’s output without the training that is meant to make the probabilities trustworthy.
jev-on-a-laptop: the interface on small local models
jev-on-a-laptop is an unofficial study that reproduces Jev-style parallel constrained decoding on stock 1.5B–8B models on an Apple Silicon laptop. It has no affiliation with TypeSafe and no access to Jev’s model. It is the most useful source here because it measures itself honestly.
Accuracy. The author rebuilt every public example case from TypeSafe’s four workflows and scored each model against TypeSafe’s own reference answers, over 343 question pairs:
| Model | Agreement |
|---|---|
| Opus (published) | 89.8% |
| DeepSeek v4.1 Flash | 89.5% |
| Sol (published) | 89.2% |
| Jev (published) | 86.6% |
| Qwen2.5-7B, local, free | 73.8% |
The free local model is about 13 points behind Jev on Jev’s own benchmark. The author also flags that an earlier five-case run had shown a tie, and calls it an artifact of the tiny sample — a useful warning about small tests in general.
Confidence. The local models’ confidence does not reliably flag errors: the 7B model was more than 0.90 confident on 13 of its 20 wrong fields. The author marks calibrated confidence as not reproduced, because a raw softmax over the candidate options is a proxy, not trained calibration. This is a finding about the reproduction, not about Jev.
Speed. Against its own naive JSON baseline, the constrained decoding was 3.4–7.9x faster locally. The author marks TypeSafe’s 40–200x claim as not applicable to a local setup.
open-jev: the interface on a diffusion model
open-jev takes a different base: typed JSON inference with DiffusionGemma, and it replays all 408 questions from Jev’s public examples. Our sources contain no accuracy or calibration result for open-jev itself, so we cannot tell you how close it gets. It also references a bundle of experiments attributed to Every; its author notes those figures are agreement with a saved model’s answers, not human-labelled accuracy, and we found no standalone publication of them.
GLiNER: the one independent head-to-head
The comparison with the strongest method is against fastino/gliner2.5-multi-v1, a zero-shot classifier you can run on a CPU. jev-benchmarks ran a pre-registered pilot on 100 held-out examples from each of three datasets, with a uniform negative control. The accuracy and coverage table is in Jev Alternatives; what that table leaves out is how firm each result is:
| Dataset | Labels | Jev minus GLiNER accuracy, 95% CI |
|---|---|---|
| AG News | 4 | +0.130 to +0.290 |
| Banking77 | 72 | +0.220 to +0.300 |
| DAIR Emotion | 6 | −0.070 to +0.150 |
On the first two, Jev’s lead is clear even at the low end of the interval. On DAIR Emotion the accuracy difference is unresolved, and calibration goes the other way: Jev’s Brier score was 0.846 against GLiNER’s 0.668, its log loss 5.588 against 1.381, and it put zero probability on the true label for 16% of examples. The authors call the result deliberately mixed.
GLiNER ran on an Apple M4 Max CPU while Jev was called over the network from France, so the latency figures (about 44 ms against 236–256 ms on the small label sets) describe that deployment, not the models.
A classifier you fine-tune yourself
We have no benchmark for this, and none appears in our sources. What we have is the trade-off. Jev cannot be trained on your data; its behaviour on your domain comes only from what you put in the request. A model you fine-tune learns your data’s quirks directly. One hands-on review put it simply: for a stable, narrow domain, a conventional small classifier may still be the simpler choice.
Three facts tilt the decision:
- Data handling. Everything in Jev’s
stategoes to TypeSafe, and zero data retention is offered only to enterprise customers. A local model sends nothing anywhere. - Language. TypeSafe says English is where Jev is most accurate; other languages, including CJK scripts, are handled but not as well. A model trained on your own non-English data may do better.
- Dependence. TypeSafe’s rate limits change without notice, it has had at least one outage from demand, and it says it cannot yet prove its pricing is not subsidised. A model you host has none of those risks, and all of the maintenance.
When to run your own model
- The data cannot leave your infrastructure.
- The label set is small and stable, and you already have plenty of labelled examples.
- Your inputs are mostly in a language other than English, and you can train on them.
- You need an answer without a network round trip.
When to use Jev
- You have few or no labels yet, or the labels change often: editing a question takes a minute, retraining does not.
- The label set is large — Jev’s clearest independent win was on a 72-label task.
- You want probabilities trained for calibration rather than a softmax proxy, and you will still check them on your own data.
A practical path is to start on Jev, log its answers and the ones people correct, and train your own model once you have a large, stable labelled set — then test which one wins on your data. You can start that log in the playground; if you are choosing how to access hosted Jev, see How to get Jev, and for general LLMs, Jev vs LLM classification.
よくある質問
- Is Jev open source?
- No. TypeSafe has not published Jev’s weights, parameter count, architecture or full training recipe. TechCrunch reports that outside observers suspect it is built on an open-weight LLM, but that is unconfirmed.
- Can I run Jev locally?
- No. Jev is only available as a hosted API. If the data cannot leave your machines, you need a different model.
- What is the best open-source alternative to Jev?
- It depends on the job. For a small label set that must run locally, a zero-shot classifier like GLiNER is the only option with an independent head-to-head against Jev. For a stable label set with plenty of labelled data, a classifier you train yourself. The projects that imitate Jev’s interface on small LLMs are good experiments, not yet replacements.
- Do the open reproductions give calibrated probabilities?
- Not the one that measured it. jev-on-a-laptop reports that its confidence is a raw softmax over candidate options rather than trained calibration, and that its 7B model was more than 0.90 confident on 13 of its 20 wrong fields.
関連
パイプラインに組み込む
新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。