Jev content moderation: classify user comments as allow, review or remove
Paste a user comment. Jev returns whether to allow it, send it to a human or remove it, and — separately — how likely it is to be harassment or off-platform spam, as probabilities your moderation queue can route on, not a paragraph to parse.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Decide whether a user comment is allowed, needs review or should be removed.
Edit freely — Jev answers the scenario's questions about whatever you put here.
What this moderator checks
The example above is a marketplace comment that breaks more than one rule at once: it calls the seller a scam, invites buyers to DM for a cheaper deal off-platform, and insults the person in the photo. One call asks three questions:
| Question | Type | What you get back |
|---|---|---|
verdict | Choice | A probability for allow, review and remove, plus a confidence |
harassment | Noul | Probability the comment insults or harasses a person |
spam | Noul | Probability the comment tries to move buyers off-platform or advertises |
The Choice carries the decision, and its options come with a short description each — "Borderline, send to a human moderator" for review — so Jev knows what each bucket means on your platform. The two Nouls say why a comment is in trouble. That is useful for logs, for appeals, and for rules that treat the two differently: you might hide spam silently but warn the author of an insult.
All three are answered in parallel against the same comment, so asking three questions costs almost no extra latency and still counts as one standard evaluation.
How to write moderation questions
Write your policy into the criteria. Jev reads questions literally; TypeSafe’s notes say scoping words, negations and implied conditions are taken at face value. "Does the comment insult or harass a person?" will not know that your community allows harsh criticism of a product but not of a seller. Say so:
{
"type": "noul",
"instructions": "Does the comment insult or harass a person?",
"criteria": {
"true": "It attacks, mocks or demeans a specific person, including the seller or anyone pictured.",
"false": "It criticises a product, price or service without attacking a person."
}
}
One Noul per rule. Harassment, spam, hate, personal data, self-promotion — each gets its own question, so a high score points to a specific policy. Up to five questions on the same comment still cost 1 credit; up to 12 are allowed, at 3 credits.
Keep the review option honest. A three-way Choice with a real "borderline" bucket gives you a probability mass you can route to people, instead of forcing every close call into allow or remove.
Keep counting in code. Rules such as "more than three links" or "posted five times in a minute" are counts. TypeSafe documents that Jev does not count reliably; compute those in code and let Jev judge meaning.
For more on phrasing, see Designing Questions Jev Can Answer.
From one check to a moderation queue
- Label a sample. Take 50 to 200 real comments your moderators have already decided and run them through the same questions.
- Pick three bands. Auto-publish comments where
allowis high and both Nouls are low, auto-hide whereremoveis high, and queue the rest for a person. Tune the edges on your labels. - Trim the state. Send the comment and only the context a moderator needs. TypeSafe notes that unrelated material in the state costs accuracy, so don’t paste the whole thread.
- Record the model version. Every response carries a
modelfield such asjev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request. - Call it from your pipeline. The same evaluation runs through the System One API with an API key, at the same credits as the playground.
At TypeSafe’s list price of $0.042 per million input tokens, a typical three-question check like this one is on the order of $0.00002 in tokens — cheap enough to screen every comment rather than a sample. The measurement is in What a Jev Evaluation Costs.
Check it before you trust it
- Measure against your moderators, not the example. Community norms differ; a threshold from someone else’s platform is not yours.
- Don’t reuse thresholds across question types. TypeSafe documents that a Noul and an equivalent Choice can disagree. Tune
verdictand each Noul separately. - Watch the review band. If most comments land in the middle, the criteria are too vague — tighten them before adding moderators.
- Test every language you serve. Accuracy is best in English.
Where Jev is the wrong tool
- You must show a written reason to the author or keep a narrative record for each removal.
- The content is not text. Jev reads a text state; images and video need a different model.
- The rule is exact. Banned words, blocked domains, rate limits and link counts belong in code.
- Spammers can probe your scores. Cheap, stable scores make it easy to test wording until something slips through, so never expose them to posters.
The same allow / review / remove pattern — a decision Choice plus one Noul per reason — carries over to support tickets and other queues; see support ticket triage and Jev vs LLM classification.
Frequently asked questions
- Can Jev replace human moderators?
- No, and it is not built to. It is built to sort the queue: publish the comments it is confident are fine, remove the clear violations, and send the uncertain middle to a person. How much it clears on its own depends on your policy and your thresholds.
- Can Jev explain why a comment was removed?
- No. Jev returns probabilities, not text. If you have to show users a reason, derive it from which question fired — harassment or spam — or keep a generative model for the notice itself.
- What does moderating one comment cost?
- On JevStation the three questions in this example cost 1 credit together for any comment under 8,000 characters. You can ask up to five questions for the same credit.
- Does it work for non-English comments?
- It handles them, but TypeSafe says English is where accuracy is best and other languages, including CJK scripts, are handled less well. Measure on your own non-English comments before relying on a threshold.
Related
- Prompt injection detectorFlag instruction overrides and data-exfiltration requests in LLM inputs and get a pass / sanitize / block decision.
- Support ticket triageRoute a support ticket to the right team, rate its urgency and flag churn risk, each with a probability.
- Jev vs LLM classificationJev against a prompted GPT- or Claude-class model for classification: cost, latency, output shape, consistency and calibration.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.