Jev Use Cases: 12 System One Model Decisions
What can Jev do? Twelve decisions it makes in an agent or pipeline, from ticket triage to tool-call guards, each with its question set and a free tool page.
Jev makes decisions whose possible answers you can list in advance: which team owns a ticket, whether a tool call matches what the user asked, whether an answer is grounded in its sources. It never writes the reply, the summary or the reason. Each use case below is one of those decisions, written as a set of typed questions (Choice, Score, Noul) that you can run on your own text on its tool page.
Every example here uses between two and four questions on a short state, so on JevStation each one is a standard evaluation: 1 credit, as long as the state stays under 8,000 characters and the call asks five questions or fewer. The questions are answered in parallel against the same state, so a third or fourth question adds almost no latency.
A quick reminder of the three shapes, because every use case is built from them:
| Type | What you ask | What comes back |
|---|---|---|
| Choice | Pick one option from a list you wrote | A probability per option, plus confidence |
| Score | Place the state on an ordered scale | A position on the scale, the distribution, confidence |
| Noul | A yes/no question | The probability that the answer is yes |
Support and operations
1. Support ticket triage
A Choice picks the owning team (billing, integrations, bug, account), a five-level Score rates urgency from "Can wait" to "Critical", and a Noul gives the probability the customer is about to churn. The useful output is often the split: a ticket that mentions Stripe but is really about connecting a service shows up as probability shared between two teams, and a low confidence tells you not to auto-assign it. Email triage at scale is one of the builds LangChain names in its write-up of Jev. Support ticket triage
2. Review sentiment
A Choice reads the overall sentiment (positive, mixed, negative), a Score maps the review to an implied star rating, and a Noul asks whether the reviewer wants a refund. The refund Noul is the one that usually triggers an action; the other two feed dashboards. Be careful here: the closest independent data point, an emotion-classification dataset in jev-benchmarks, is where Jev did worst (accuracy 0.48, and poorly calibrated). Sentiment is not emotion, but measure before you trust it. Review sentiment
Safety and moderation
3. Content moderation
A three-way Choice decides allow, review or remove, and one Noul per rule (harassment, off-platform spam) says why a comment is in trouble. The review option is deliberate: it gives the uncertain middle somewhere to go other than a forced yes or no, and that middle is what you route to human moderators. Counts such as "more than three links" stay in code, because TypeSafe documents that Jev does not count reliably. Content moderation
4. Prompt injection detection
Two Nouls separate two harms: an attempt to override or extract the assistant's instructions, and a request to send data to an outside destination. A Choice turns both into pass, sanitize or block before the input reaches your LLM. Jev has no instruction channel to hijack, but TypeSafe's own jaggedness notes say adversarial content can still move the answer, so treat it as one layer, not a security boundary. Prompt injection detector
5. Tool-call guard
The state holds the user's request and the tool call an agent proposes. A Noul asks whether the call matches what the user asked, a Score rates the risk from "Harmless" to "Destructive", and a Choice decides execute, confirm or block. The example is an agent asked to clean staging data that proposes DELETE FROM orders; on production. LangChain's experimental AutoModeMiddleware applies the same idea to any agent, and a Vercel engineer told TechCrunch the company replaced an LLM classifier that reviewed commands for safety with Jev, and got results five to 18 times faster. Tool-call guard
Sales and hiring
6. Lead qualification
A Score rates how well an inbound lead fits a B2B product, a Choice places it in a buying stage (researching, evaluating, ready to buy), and a Noul asks whether it mentions a concrete timeline. That is enough to decide who gets a same-day reply. The one public practitioner report on classifying business emails, from Bryo AI's CTO via TechCrunch, found Gemini slightly more accurate but 10 to 20 times more expensive; what interested the CTO most was that Jev hands back a real probability. Lead qualification
7. Resume screening
A Score rates the match between a candidate and a role, a Noul asks whether they meet the seniority bar, and a Choice suggests reject, phone screen or onsite. Recruiting is on TypeSafe's own use-case map. This is also the use case where the lack of a written reason matters most: an analysis in 36Kr argues that decisions without an explanation often hit a compliance wall in regulated fields. Use it to order a queue for a person, not to reject people on its own. Resume screening
Agents and RAG
8. LLM answer judge
A Noul checks whether an answer is consistent with reference facts, a Score grades overall quality, and a Choice decides ship, revise or reject. This is the use case with the strongest independent evidence: in an AI-control pilot, a single Noul ranked backdoored code above honest code with AUROC 0.976. The same pilot found clear failure modes, which Jev as a Judge covers in full. Jev as a judge
9. RAG evaluation
Two Nouls ask whether every claim in the answer is supported by the retrieved context, and whether the context contained what the question needed at all. A Score rates the answer, and a Choice names where a wrong answer failed: retrieval (the context lacked it) or generation (the context had it and the answer misstated it). That last question is what turns an evaluation into a fix, because the two failures have different owners. RAG evaluation
10. Citation check
The state holds a source passage and a claim. A Choice decides supported, partially supported or unsupported, and a Noul asks whether the claim overstates the source. The example claims "most IT leaders" planned to spend more when the survey says 38%, which is the kind of quiet overstatement worth catching. The pattern pairs well with an LLM, as beam.ai suggests: let the LLM extract, then hand the extracted value and its source to Jev to verify. Deciding whether 38% counts as "most" leans on numbers, which TypeSafe lists as a weakness, so test examples like it. Citation check
11. LLM router
A Score rates how demanding a prompt is, a Choice picks a small, medium or frontier model tier, and a Noul asks whether the answer needs code. An independent Classmethod test got the expected tier on all 40 calls, with one prompt per tier, and its author says plainly that it is not an accuracy benchmark. Jev picks the tier; your code sends the request. LLM router
12. Context compaction
A Noul asks whether a stretch of transcript holds a decision or constraint the agent must remember, and a Score rates how important it is to keep. Instead of summarising, compaction deletes what scores low and keeps the rest word for word. fast-jev-compaction is a worked example: a Claude Code plugin that asks two yes/no questions per tool call and falls back to the built-in summary when the saving is small. Context compaction
When Jev is not the right tool
The twelve decisions above share a shape. When your problem does not have it, a different tool is better:
- You need text. Replies, summaries, critiques and written reasons need a generative model. TypeSafe's notes say forcing Jev to generate text "will not work well".
- You cannot list the answers. Open-ended questions have no Choice to ask.
- The decision is counting, arithmetic or dates. "Three tickets this week" or "past the renewal date" belong in code; pass the result into the state.
- The rule is exact. Blocked domains, keyword lists and size limits are code, not classification.
- Most of your input is not English. TypeSafe says other languages, including CJK scripts, are handled but not as well.
- Whoever is being judged can probe the scores. In the AI-control pilot, an attacker who could check the score before submitting got about a third of backdoors through instead of a tenth.
- Your label set is stable and you have lots of labelled data. A classifier you train yourself may be the simpler, cheaper choice at high volume. Jev Alternatives walks through that trade.
How to pick your first use case
Start where three things are already true, a test beam.ai puts well: you know the set of possible answers, you make the same decision many times, and you can act on a confidence score. In practice that points to a queue a person already sorts by hand, such as tickets, comments or leads, because the past decisions are your labelled data.
- Open the tool page closest to your queue and run the example.
- Replace the state with twenty real inputs whose right answer you know, in the playground. With more rows, the batch page runs one question set over every row of a CSV.
- Rewrite the options in your own words, with a one-line description each.
- Look at where the confidence separates right answers from wrong ones, and set three bands: act, review, reject.
The first 200 evaluations are free, which covers that test several times over. What a System One Model Is explains why the three conditions matter, and Designing Questions Jev Can Answer covers the wording. To browse all twelve pages in one place, start at tools.
Sources
- Question sets and example states. The 12 JevStation tool pages linked above (our configuration)
- Credit tiers and the free grant. Pricing (JevStation)
- Email triage at scale;
AutoModeMiddleware. LangChain (third party) - Vercel command-safety review; Bryo AI email classification. TechCrunch via Yahoo (third party, practitioner reports without published data)
- Recruiting and model routing as use cases. TypeSafe use-case map (vendor)
- Counting, dates, adversarial content, generation. Jev 1.13 jaggedness (vendor)
- Language support. TypeSafe model docs (vendor)
- Three-conditions heuristic; extract-then-verify pattern. beam.ai (third party)
- Emotion-classification result. jev-benchmarks (independent)
- Backdoor-monitor pilot, probing attacker. LessWrong (independent)
- Routing test. DevelopersIO (Classmethod) (independent)
- Compliance concern. 36Kr (third party)