Jev as a judge: grade LLM answers with typed verdicts
Paste a question, the reference facts and a model’s answer. Jev returns whether the answer is consistent with the facts, how good it is on a five-level rubric, and whether to ship it — as probabilities your code can threshold, not a paragraph you have to parse.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Grade a model's answer against reference facts and decide whether to ship it.
Edit freely — Jev answers the scenario's questions about whatever you put here.
What this judge checks
The example above grades a support-bot answer against the knowledge-base snippet it should have used. It asks three questions in one call:
| Question | Type | What you get back |
|---|---|---|
correct | Noul | Probability the answer is consistent with reference |
quality | Score | A position on Wrong → Excellent, plus confidence |
verdict | Choice | A probability for ship, revise and reject |
Three questions still cost one credit, and they are answered in parallel against the same state, so asking all three adds almost no latency. The answer in the example contradicts the reference — it says the free plan has no API access — so the interesting output is how strongly correct leans to "no" and whether verdict picks reject or revise.
How to write judge questions that work
One failure mode per Noul. "Is this answer good?" mixes accuracy, tone and completeness into one probability you cannot act on. "Does the answer contradict the reference?" has one meaning. Add a second Noul for each other failure you care about, such as "Does the answer promise something the policy does not allow?"
Give the Noul criteria. A true and false description tells Jev where the line is:
{
"type": "noul",
"instructions": "Does `answer` state or imply something that contradicts `reference`?",
"criteria": {
"true": "The answer asserts a fact the reference contradicts.",
"false": "Everything the answer claims is supported by, or absent from, the reference."
}
}
Use a Score when severity matters. A five-level rubric with a short description per level ranks answers better than a yes/no when most outputs are "partly right".
Put the decision in a Choice. Asking directly for ship / revise / reject gives you a distribution over actions, which is easier to route on than combining three numbers yourself. Keep the other two questions for logging and threshold tuning.
For more on phrasing, see Designing Questions Jev Can Answer.
From one check to a grading step
- Collect labelled examples. 50 to 200 real answers, each marked good or bad by a person, are enough to see where Jev’s probabilities separate them.
- Pick three bands, not one threshold. Clear the answers Jev is confident about, escalate the middle band to a stronger judge or a human, and block the rest. In the public pilot this pattern let Jev clear about 80% of the queue on its own.
- Record the model version. Every response carries a
modelfield such asjev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request. - Call it from your pipeline. The same evaluation runs through the System One API with an API key, and charges the same credits as the playground.
Check it before you trust it
- Compare against your labels, not the example. A threshold that works on someone else’s data is not your threshold.
- Test the wording. A Noul and its negation, or the same idea as a Noul versus a Choice, can give different probabilities. Tune the thresholds on the exact questions you ship.
- Watch the middle band. If most answers land between your thresholds, the questions are too vague — split them.
Where Jev is the wrong judge
- You need a written reason for a reviewer, an audit trail or feedback to the model being graded.
- The judgement involves counting, arithmetic or dates. TypeSafe documents these as weaknesses.
- The judged system can query the judge. Cheap, stable scores make it easy to probe for what slips through.
The evidence behind these limits — and the pilot numbers quoted on this page — is laid out with sources in Jev as a Judge: What One Yes/No Question Catches, and What It Misses.
Frequently asked questions
- Can Jev replace an LLM-as-a-judge?
- For verdicts you can enumerate in advance — correct or not, pass or fail, a rubric level — yes. It does not write a critique or a reason, so keep a generative judge where you need an explanation or open-ended feedback.
- How accurate is Jev as a judge?
- The best public test is an independent AI-control pilot where one Jev yes/no question ranked backdoored code above honest code with AUROC 0.976. Accuracy on your task depends on your questions, so measure it on your own labelled examples before relying on a threshold.
- What does one judgement cost?
- On JevStation a standard evaluation costs 1 credit, whether it has one question or five. At Jev’s list price the tokens for a typical check cost about $0.00002.
- Can the answer being judged manipulate the verdict?
- Jev has no instruction channel — text in the state is read as content — and in the public pilot notes addressed to the judge did not rescue bad outputs. But if the judged system can call the judge repeatedly and keep the best-scoring attempt, it can learn to slip past, so never expose judge scores to it.
Related
- Prompt injection detectorFlag instruction overrides and data-exfiltration requests in LLM inputs and get a pass / sanitize / block decision.
- LLM routerClassify a prompt’s difficulty and pick a small, medium or frontier model before you spend on the completion.
- Jev vs LLM classificationJev against a prompted GPT- or Claude-class model for classification: cost, latency, output shape, consistency and calibration.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.