Designing Questions Jev Can Answer: Choice, Score and Noul
The three Jev question types, Choice, Score and Noul, decide what your code can branch on. What each returns, their limits, and how to write atomic questions, sharpen rubrics and gate agent actions on confidence.

A Jev question set is not a prompt. It is the schema of a decision: the answer comes back typed, and its shape is the shape of the branch your code can take. Write the questions carelessly and a confident model hands you a useless number. Write them well and the control flow falls out of the response.
This is the part of working with System One models where the feedback loop is shortest and the effect on quality is largest — and it is mostly mechanical once you know the rules.
The call you are designing for
With the LangChain integration, a question set is ordinary Python — the questions are declared once on the classifier, and every call evaluates the same set against a new state:
from langchain_typesafe import Choice, Noul, Score, TypeSafeClassifier
classifier = TypeSafeClassifier(
questions={
"urgent": Noul(instructions="Does this need attention right now?"),
"team": Choice(
instructions="Which team should pick this up?",
criteria={
"infra": "Deploys, availability, and on-call incidents.",
"billing": "Payments, invoices, and subscriptions.",
},
),
}
)
response = classifier.invoke(
"The deploy failed twice and customers are seeing 500s. Can someone look now?"
)
urgency = response.nouls["urgent"].noul
owner, certainty = response.choices["team"].choice, response.choices["team"].confidence
Without a framework it is a single POST to https://api.typesafe.ai/v1/systemone, and the body is what you would expect: a state, a model, and a map of questions. The state can be a string, structured data, or a list of messages; the questions are the part you actually author.
The three Jev question types at a glance
Every question is one of three types. Pick the type from the shape of the branch you want, not from how the question reads in English:
| Type | Use it when | Returns | Limits |
|---|---|---|---|
| Noul | The branch is yes or no | noul: the probability the answer is yes, 0 to 1 | No separate confidence; criteria optional |
| Choice | The branch is one of several unordered options | choice, a probability for every option (sum to 1), confidence | Up to 255 options |
| Score | The branch depends on how much, on a rubric | score, per-level probabilities, legend, confidence | 2 to 10 ordered levels |
Three details from TypeSafe's primitives docs that change how you write them:
- A Noul is not a degree. 0.6 means "probably yes", not "somewhat". If you want a degree, ask a Score.
- A Score is a weighted mean. It runs from 0 to the number of levels minus 1 and can land between levels, so two different distributions can produce the same score. Read the distribution when it matters.
- Questions in one request are independent. One answer is never context for another, so a question cannot refer to the answer of the one before it.
The rest of this guide is about writing each of them well.
Rule 1: one question, one judgement
The classic mistake is the compound question: is this ticket urgent and should we auto-resolve it? Two factors, one number, and no way to act on the result. Ask each factor separately and combine the answers in your own code.
{
"intent": {
"type": "choice",
"instructions": "What is the customer asking for?",
"criteria": {
"billing": "Invoices, charges, refunds, or payment failures",
"bug": "Something is broken or behaving incorrectly",
"how_to": "A question about using the product as intended",
"other": "None of the above"
}
},
"needs_human": {
"type": "noul",
"instructions": "Does this require a human decision before a reply is sent?"
},
"severity": {
"type": "score",
"instructions": "How severe is the customer's problem?",
"criteria": [
"Cosmetic",
"Annoying but workable",
"Blocking one workflow",
"Blocking all work",
"Data loss or outage"
]
}
}
The routing rule then lives in code, where you can read it, test it, and retune it:
const route =
intent === 'billing' && !needsHuman && severity <= 2 ? 'auto-reply' : 'human';
Because the questions are evaluated in isolation, a rubric you got wrong on severity does not contaminate intent. You can fix one question without re-validating the whole set.
Rule 2: instructions are the job, criteria are the rubric
Every question carries instructions — what to judge — and its own criteria: an option map for Choice, an ordered level array for Score, and optional true/false descriptions for Noul.
The criteria are where quality is won or lost, because they define the boundary between neighbouring labels. "High" and "medium" mean nothing on their own; "blocking one workflow" versus "blocking all work" is a distinction a model can apply consistently.
Vague criteria produce flat distributions, and a flat distribution is the model telling you it cannot separate your options. Sharpen it empirically: run the set over examples you have already labelled, look at where probability mass leaks between two adjacent options, and rewrite the descriptions that failed to do the separating.
One anti-pattern worth naming: do not smuggle the answer into the instructions. "Is this a billing issue? Billing issues mention invoices, refunds, charges" is a keyword filter wearing a model costume, and it will fail on the first customer who writes "you charged me twice" without using any of those words. State the judgement; let the criteria carry the boundary; let the state carry the evidence.
Rule 3: the state is evidence, not instruction
Send the record, not a description of the record. A JSON object with the fields your system already has — plan, tenure, error code, ticket body — gives the model something to judge against, and structured state costs the same as the prose version you would otherwise write.
Keep the question set stable across a batch and vary only the state. That is what makes a set reusable: same questions, new evidence, comparable answers you can chart over time.
Rule 4: read the distribution, not just the winner
Every answer has a winner and a shape. Choice returns choice, probabilities, and confidence. Score returns a score that can land between levels, a legend mapping each level back to its description, the distribution, and confidence. Noul returns a single probability and nothing else.
When the top two options are close, the model is telling you the case is genuinely ambiguous — that is information about the case, not only about the model. A three-band policy turns that into behaviour:
const BANDS = { auto: 0.9, review: 0.6 };
function decide(choice: string, confidence: number) {
if (confidence >= BANDS.auto) return { act: choice };
if (confidence >= BANDS.review) return { act: 'confirm', proposed: choice };
return { act: 'escalate' };
}
Noul has no confidence value, so threshold the probability directly: above 0.98 treat the statement as true, between roughly 0.7 and 0.98 route it to review, below that treat it as unknown and ask — of the user, not of the model.
Set the boundaries by consequence. Auto-sending a reply and deleting a branch are not the same bet, and the thresholds should say so.
Rule 5: keep the set small
A request is billed on the state plus every question, which makes the question set the one part of your cost you fully control. Output tokens, helpfully, are free — there is no output to generate. Adding questions barely moves latency, because they run in parallel; what it moves is input tokens, and duplicate questions buy you nothing except a second opinion you did not ask for.
There is a hard ceiling too, and it is shared with the state: the model's documented 64k request budget covers the state plus all of your questions combined, with a 32k budget for the state plus the single longest question. A question set you keep growing eventually competes with the evidence it is supposed to judge.
JevStation's question editor enforces its own limits for you: up to 12 questions per set, 12 options per Choice question, 2 to 12 levels per Score question, and 600 characters of instructions per question. Each response also reports usage.input_tokens and usage.output_tokens, so you can watch a set's cost as you grow it.
The rule of thumb is four to six sharp questions. If you find yourself writing twelve, some of them are factors that belong in a formula instead.
The economics make this easy to accept: at the published Jev 1.13 rate of $42 per billion input tokens, a thousand-token ticket with six questions on top costs well under a hundredth of a cent. Spend the budget on sharper questions rather than more of them.
What Jev is bad at
TypeSafe publishes the jagged edges of the current model, and designing around them is part of the job:
- It reads literally. It answers the question you wrote, not the one you meant, so state the exact condition in the instructions and describe each option's boundary in the criteria.
- Do the maths in code. Counting, numeric comparison, and date ordering are weak spots, and the error grows with the size of the thing being counted. Ask Jev for a judgement, then compute with the answer.
- Context rots. Unrelated material in the state costs accuracy, so retrieve and trim before you send, rather than pasting the whole record and hoping.
- Question types do not agree with each other. A Noul and an equivalent yes/no Choice on the same state can return different answers, and a question plus its negation need not sum to 1. Do not carry a threshold tuned on a Noul over to a Choice, or hold separate questions to arithmetic identities.
- Score is a threshold check, not a measurement. The level descriptions are weakly calibrated as magnitudes, so compare a score against a cutoff instead of interpolating an exact value from it.
- English is the strongest language. Other languages, Chinese included, are handled but not equally well — test on your own content before you rely on Jev for a non-English workload.
Putting it together
A support triage set — intent, severity, and whether a human is needed — becomes a routing function you can explain to a support lead in one screen:
type Triage = {
intent: { choice: string; confidence: number };
needs_human: { noul: number };
severity: { score: number; confidence: number };
};
function route(a: Triage) {
if (a.needs_human.noul >= 0.98) return 'human';
if (a.severity.score >= 4 && a.severity.confidence >= 0.8)
return 'page-oncall';
if (a.intent.choice === 'how_to' && a.intent.confidence >= 0.9)
return 'auto-reply';
return 'queue';
}
Then treat the set like a test suite. Freeze a handful of cases with the labels you would assign by hand, re-run them whenever you touch instructions or criteria, and compare distributions rather than winners. It is cheap — one call per case, a few questions each — and it is the only way to tell a better question set from a luckier one.
Takeaways
- One judgement per question, combined in code: coefficients are easier to retune than prompts.
- Instructions describe the job, criteria describe boundaries, the state carries the evidence.
- Confidence is a policy input, not a decoration: act high, confirm in the middle, escalate low — and set the boundaries by what a mistake costs.
The fastest way to internalise this is to break a decision you actually own into three questions and run it in the playground — you get the distributions immediately, and the answer shapes tell you which question needs work. The docs cover the request and answer fields in full, and Building a Harness with Jev covers where these calls belong in an agent loop.