RAG evaluation with Jev: grounded answers and where the pipeline failed
Paste a user question, the context your retriever returned and the answer your model wrote. Jev returns whether the answer is grounded in the context, whether the context held what was needed, how well the question was answered, and — if something went wrong — whether retrieval or generation is to blame.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Check whether a RAG answer is grounded in the retrieved context, and whether retrieval or generation failed.
Edit freely — Jev answers the scenario's questions about whatever you put here.
What this RAG evaluation checks
The example above is a refund question. The retrieved context says refunds are issued within 2 business days of approval and that card issuers usually post them 5–10 business days after that. The answer says refunds are instant and arrive the same day. Jev answers four questions about it in one call:
| Question | Type | What you get back |
|---|---|---|
grounded | Noul | Probability that every claim in the answer is supported by retrieved_context |
context_relevant | Noul | Probability that the context contains the information needed to answer |
answer_quality | Score | A position on Wrong → Poor → Partial → Good → Complete, plus confidence |
failure | Choice | A probability for none, retrieval and generation, plus confidence |
This example is useful because the two stages disagree. The retriever did its job — the timing is right there in the context — and the model then contradicted it. So the outputs worth reading are how low grounded goes, how high context_relevant stays, and whether failure leans towards generation rather than retrieval. A single "is this answer good?" score would flag the answer but not tell you which half of the pipeline to fix.
Four questions on a short state are a standard evaluation, 1 credit on JevStation. The questions are answered in parallel against the same state, so asking about both stages adds almost no latency.
How to write RAG questions that work
Keep groundedness and correctness apart. An answer can be grounded in a stale document and still be wrong for the user, or correct from general knowledge and not grounded at all. grounded asks only about support in the context. If you also need correctness against a reference, add a separate question — see Jev as a judge.
Ask about the context on its own. context_relevant is the retrieval check. When it is low, the generator had nothing to work with, and a better prompt will not help. Give it criteria so Jev knows where the line is:
{
"type": "noul",
"instructions": "Does `retrieved_context` contain the information needed to answer `question`?",
"criteria": {
"true": "A careful reader could answer the question using only the context.",
"false": "The context is off-topic, incomplete or silent on what the question asks."
}
}
Describe every failure option. The failure Choice works because each option has a one-line description. Jev reads literally, so vague options such as "bad retrieval" leave it to guess your boundaries.
Do not expect the questions to agree arithmetically. TypeSafe notes that separate questions are independent and need not add up — a Noul and an equivalent Choice can give different numbers. Read failure alongside context_relevant rather than deriving one from the other.
For more on phrasing, see Designing Questions Jev Can Answer.
From one check to an evaluation step
- Build a labelled set. 50 to 200 real question–context–answer triples, each marked by a person as grounded or not and, where wrong, which stage failed.
- Run the set in one go. Put the triples in a CSV and run the same question set over every row with batch evaluation. Compare the probabilities to your labels.
- Pick bands, not one threshold. Pass answers that are clearly grounded, regenerate or escalate the middle band, and block the rest.
- Record the model version. Every response carries a
modelfield. Log it with each result and re-check your thresholds when it changes — JevStation pins the model at deployment, so you do not choose it per request. - Call it from your pipeline. The same evaluation runs through the System One API with an API key and charges the same credits as the playground.
Check it before you trust it
- Measure against your own labels. A threshold that separates good from bad on the refund example is not a threshold for your corpus.
- Watch for injected text in retrieved documents. TypeSafe’s notes say content written to steer the model can move the answer. If your corpus includes untrusted pages, screen them with the prompt injection detector as well.
- Test non-English content separately. TypeSafe documents English as the language where accuracy is best.
Where Jev is the wrong RAG evaluator
- The answer turns on numbers or dates. Whether "within 7 days" matches "5–10 business days" is arithmetic, and TypeSafe lists math and date comparison as weaknesses. Extract and compare those values in code.
- You pass the whole corpus as context. TypeSafe warns that unrelated material in the state costs accuracy. Send only the passages the retriever returned; JevStation caps the state at 24,000 characters.
- You need to know which sentence is unsupported. Split the answer into claims and run the citation check on each one against its passage.
- You need a written critique for a reviewer or as feedback to the generator.
The evidence behind grading with a single yes/no question — and where it breaks — is laid out in Jev as a Judge: What One Yes/No Question Catches, and What It Misses.
Frequently asked questions
- How do I evaluate RAG answers for groundedness with Jev?
- Send the question, the retrieved context and the answer as one state, and ask a Noul such as "Is every claim in the answer supported by the retrieved context?". Jev returns the probability that it is, which you can threshold per answer.
- How does Jev tell a retrieval failure from a generation failure?
- Ask it directly. The example uses a Choice with three options — none, retrieval and generation — each described in one line, plus a separate Noul on whether the context contained the needed information. Log both, and fix the stage the answers point to.
- What does one RAG check cost?
- On JevStation, up to five questions on a state of up to 8,000 characters is a standard evaluation and costs 1 credit, so the four questions in the example cost 1 credit. Longer context or more than five questions costs 3 credits. A failed evaluation is refunded.
- Can Jev check each claim in a long answer separately?
- One Noul gives one probability for the whole answer. If you need to know which sentence is unsupported, split the answer into claims in code and check each one against its passage, as the citation check tool does.
- Does Jev explain why an answer is ungrounded?
- No. Jev returns typed answers — probabilities and confidence — not written reasons. Keep a generative model for the cases where a reviewer needs an explanation.
Related
- Citation checkVerify one claim against one source: supported, partially supported or unsupported, and whether it is overstated.
- Jev as a judgeCheck an LLM answer against reference facts and get a ship / revise / reject verdict with probabilities.
- Context compactionScore transcript excerpts for keep or drop, so compaction deletes what is stale and keeps the rest verbatim.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.