Citation check with Jev: is the claim supported by its source?

Paste a source passage and a claim that cites it. Jev returns whether the source supports the claim fully, partially or not at all, and how likely it is that the claim overstates what the source says — as probabilities your review process can sort on, not a paragraph to parse.

自分のテキストで試す

例を編集して実行してください。アカウントは不要で、1 日 3 回まで無料で実行できます。

シナリオ

主張が出典によって実際に裏付けられているかを確認します。

244 / 2,000

自由に編集できます。ここに入力した内容について、Jev がシナリオの質問に回答します。

What this citation check tests

The example above is a claim that looks close to its source. The source says a 2023 survey of 1,200 IT leaders found 38% planned to increase spending on observability tools and 12% planned cuts. The claim says most IT leaders planned to increase observability spending. Jev answers two questions about it in one call:

QuestionTypeWhat you get back
supportedChoiceA probability for supported, partially and unsupported, plus confidence
overstatedNoulThe probability that the claim overstates what the source says, from 0 to 1

The claim gets the direction right — more leaders planned increases than cuts — but "most" is a stronger word than 38% supports. This is the typical failure in research summaries and write-ups: not a fabrication, but a finding stretched a little in the retelling. So the outputs to read are how the supported probability splits between partially and unsupported, and how high overstated goes. A plain supported / not-supported answer would hide that difference.

Two questions on a short state are a standard evaluation: 1 credit on JevStation.

How to write claim-check questions that work

One claim per state. If the claim field holds three sentences, a single probability has to cover all of them. Split compound claims in code and check each one against the passage it cites.

Send the passage, not the paper. Pass the paragraph the citation points to. TypeSafe notes that unrelated material in the state costs accuracy, so trimming the source is part of the check.

Describe each level of support. The supported Choice defines partially as "The source supports part of the claim only". Jev reads literally, so if your editorial standard treats a hedged claim as supported, say so in the criteria.

Give overstatement its own criteria. A Noul with a true and false line makes the boundary explicit:

{
  "type": "noul",
  "instructions": "Does `claim` overstate what `source` says?",
  "criteria": {
    "true": "The claim uses stronger wording, a larger scope or more certainty than the source.",
    "false": "The claim says the same as the source, or less."
  }
}

For more on phrasing, see Designing Questions Jev Can Answer.

From one claim to a review step

  1. Extract claim–source pairs. For each citation in a draft, pair the sentence with the passage it cites. That pairing is ordinary code, not a Jev question.
  2. Run them together. Put the pairs in a CSV and run the same two questions over every row with batch evaluation.
  3. Sort, do not auto-reject. Pass claims that are clearly supported and not overstated, and send the rest to an editor with the probabilities attached so the most doubtful are read first.
  4. Record the model version. Every response carries a model field. Log it with each result and re-check your thresholds when it changes — JevStation pins the model at deployment, so you do not choose it per request.
  5. Call it from your tools. The same evaluation runs through the System One API with an API key and charges the same credits as the playground.

Check it before you trust it

  • Use your own verified claims. Collect 50 to 200 claims a person has already checked and see where Jev’s probabilities separate them.
  • Test the wording. TypeSafe’s notes show that a Noul and an equivalent Choice can disagree. Do not carry a threshold tuned on overstated over to a Choice asking the same thing.
  • Keep a person on the final call. Calibration is measured across groups of answers, not per answer, so treat each result as a priority for review rather than a verdict.

Where Jev is the wrong claim checker

  • The claim turns on arithmetic. Whether "most" fairly describes 38% is a question about wording; whether 38% of 1,200 is "over 450" is arithmetic, which TypeSafe lists as a weakness. Extract figures and compare them in code, then pass the result into the state.
  • The claim compares dates. "Before the policy changed" against a dated source is a date comparison, another documented weakness.
  • The source is very long. JevStation caps the state at 24,000 characters. For a full report, find the relevant passage first.
  • You are checking a RAG answer end to end. Use RAG evaluation, which also asks whether retrieval or generation failed.
  • You need a written explanation for the author or the reader.

The evidence behind these limits is set out in Designing Questions Jev Can Answer, and Jev as a Judge covers how a single yes/no check holds up on a real grading task.

よくある質問

How do I verify a claim against its source with Jev?
Send the source passage and the claim as two fields of one state. Ask a Choice with supported, partially and unsupported, and a Noul for whether the claim overstates the source. Both come back in the same call, with probabilities.
How is this different from the RAG evaluation tool?
RAG evaluation checks a whole answer against retrieved context and asks which stage of the pipeline failed. The citation check is narrower: one claim against one source, the check an editor or fact-checker runs on a single citation.
What does one claim check cost?
On JevStation a source of up to 8,000 characters with up to five questions is a standard evaluation and costs 1 credit, so the two questions in the example cost 1 credit. Longer sources cost 3 credits, up to 24,000 characters.
Can Jev check claims about percentages and figures?
With care. TypeSafe lists numeric precision and arithmetic as weaknesses. Jev can judge whether a claim goes beyond its source, but when the check turns on comparing numbers, extract them and compare in code.

関連

パイプラインに組み込む

新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。