Jev as a judge: grade LLM answers with typed verdicts

Paste a question, the reference facts and a model’s answer. Jev returns whether the answer is consistent with the facts, how good it is on a five-level rubric, and whether to ship it — as probabilities your code can threshold, not a paragraph you have to parse.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Grade a model's answer against reference facts and decide whether to ship it.

289 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

What this judge checks

The example above grades a support-bot answer against the knowledge-base snippet it should have used. It asks three questions in one call:

QuestionTypeWhat you get back
correctNoulProbability the answer is consistent with reference
qualityScoreA position on Wrong → Excellent, plus confidence
verdictChoiceA probability for ship, revise and reject

Three questions still cost one credit, and they are answered in parallel against the same state, so asking all three adds almost no latency. The answer in the example contradicts the reference — it says the free plan has no API access — so the interesting output is how strongly correct leans to "no" and whether verdict picks reject or revise.

How to write judge questions that work

One failure mode per Noul. "Is this answer good?" mixes accuracy, tone and completeness into one probability you cannot act on. "Does the answer contradict the reference?" has one meaning. Add a second Noul for each other failure you care about, such as "Does the answer promise something the policy does not allow?"

Give the Noul criteria. A true and false description tells Jev where the line is:

{
  "type": "noul",
  "instructions": "Does `answer` state or imply something that contradicts `reference`?",
  "criteria": {
    "true": "The answer asserts a fact the reference contradicts.",
    "false": "Everything the answer claims is supported by, or absent from, the reference."
  }
}

Use a Score when severity matters. A five-level rubric with a short description per level ranks answers better than a yes/no when most outputs are "partly right".

Put the decision in a Choice. Asking directly for ship / revise / reject gives you a distribution over actions, which is easier to route on than combining three numbers yourself. Keep the other two questions for logging and threshold tuning.

For more on phrasing, see Designing Questions Jev Can Answer.

From one check to a grading step

  1. Collect labelled examples. 50 to 200 real answers, each marked good or bad by a person, are enough to see where Jev’s probabilities separate them.
  2. Pick three bands, not one threshold. Clear the answers Jev is confident about, escalate the middle band to a stronger judge or a human, and block the rest. In the public pilot this pattern let Jev clear about 80% of the queue on its own.
  3. Record the model version. Every response carries a model field such as jev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request.
  4. Call it from your pipeline. The same evaluation runs through the System One API with an API key, and charges the same credits as the playground.

Check it before you trust it

  • Compare against your labels, not the example. A threshold that works on someone else’s data is not your threshold.
  • Test the wording. A Noul and its negation, or the same idea as a Noul versus a Choice, can give different probabilities. Tune the thresholds on the exact questions you ship.
  • Watch the middle band. If most answers land between your thresholds, the questions are too vague — split them.

Where Jev is the wrong judge

  • You need a written reason for a reviewer, an audit trail or feedback to the model being graded.
  • The judgement involves counting, arithmetic or dates. TypeSafe documents these as weaknesses.
  • The judged system can query the judge. Cheap, stable scores make it easy to probe for what slips through.

The evidence behind these limits — and the pilot numbers quoted on this page — is laid out with sources in Jev as a Judge: What One Yes/No Question Catches, and What It Misses.

Frequently asked questions

Can Jev replace an LLM-as-a-judge?
For verdicts you can enumerate in advance — correct or not, pass or fail, a rubric level — yes. It does not write a critique or a reason, so keep a generative judge where you need an explanation or open-ended feedback.
How accurate is Jev as a judge?
The best public test is an independent AI-control pilot where one Jev yes/no question ranked backdoored code above honest code with AUROC 0.976. Accuracy on your task depends on your questions, so measure it on your own labelled examples before relying on a threshold.
What does one judgement cost?
On JevStation a standard evaluation costs 1 credit, whether it has one question or five. At Jev’s list price the tokens for a typical check cost about $0.00002.
Can the answer being judged manipulate the verdict?
Jev has no instruction channel — text in the state is read as content — and in the public pilot notes addressed to the judge did not rescue bad outputs. But if the judged system can call the judge repeatedly and keep the best-scoring attempt, it can learn to slip past, so never expose judge scores to it.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.