Blog — Guides and Engineering Notes
Product updates, engineering notes and practical guides on building with Jev, typed questions and structured answers your code can branch on.
Latest posts

Jev as a Judge: What One Yes/No Question Catches, and What It Misses
Can a model that never writes text be your LLM judge? The best public test says one Jev question ranks backdoored code at AUROC 0.976 for about $0.04 per thousand checks. Here is where that holds, and where it breaks.

Jev Alternatives: When an LLM, GLiNER or Your Own Classifier Is the Better Tool
Jev turns text into typed decisions with probabilities. So do structured-output LLMs, zero-shot classifiers like GLiNER, and models you train yourself. What each one is better at, with the measured numbers where they exist.

fast-jev-compaction: Compaction Without the Summary
A Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored, stale ones are dropped, and everything kept stays verbatim.

What a System One Model Is, and What It Is Not
System One models return typed decisions and calibrated probabilities instead of generated text. What that means, how RLCD differs from RLHF, and where the independent evidence is still thin.

What One Jev Evaluation Costs, Measured
We made 11 real Jev calls across every state size and question count the app allows, fitted the token model, and measured the cost: $0.000012 to $0.000377 per evaluation.

Building a Harness with Jev
An agent loop is only as fast as its slowest decision. Where Jev, the System One model from TypeSafe, fits: model routing, action gating, escalation.

Designing Questions Jev Can Answer: Choice, Score and Noul
The three Jev question types, Choice, Score and Noul, decide what your code can branch on. What each returns, their limits, and how to write atomic questions, sharpen rubrics and gate agent actions on confidence.