/docs
Using Jev
Jev evaluates a state against typed questions and returns structured answers: typed values and probability distributions your code can branch on. Here is how to run one.
Overview
Jev is TypeSafe's flagship model and the first System One model; JevStation is the workspace built around it. Large language models are built to produce text for people to read, so asking one for a judgement your code will consume means coercing text generation into structured output and parsing the result back. Jev works the other way round: send a state plus typed questions and it returns typed answers directly. No text generation, no parsing, no prose to interpret.
Quick start
Run an evaluation in the playground in three steps, then call the same model from your own code.
Paste the state
Open the playground and paste whatever Jev should judge: a support ticket, a review, a spec, or a JSON record. The playground calls the same contract your code will.
Ask your questions
Add one question per judgement and mix Choice, Score, and Noul in the same call. Every question is evaluated against the same state.
Read the answers
Each answer comes back under the id you chose, with its probability distribution and, for Choice and Score, a confidence value.
Call it from code
The playground is the same contract your code will call — design a question set there, then call it from your own backend. That endpoint is `POST /api/v1/systemone`, authenticated with an API key you create in Settings (`Authorization: Bearer sk_…`), and it consumes the same credits a playground run does. If you would rather talk to the model directly, without JevStation metering, TypeSafe's own API is the supported path.
Question types
TypeSafe exposes three AI primitives. Each asks a different kind of question and returns a different kind of answer.
| Question type | Goal | Returns |
|---|---|---|
| Choice | Choose one option from a list you define | choice, probabilities, confidence |
| Score | Rate the state against an ordered rubric | score, probabilities, confidence |
| Noul | Is this statement true? | noul (0–1) |
Atomic questions, composed in code
System One models work best when each question asks one specific, well-scoped thing: the kind of judgement a knowledgeable person could make in a few seconds given the right context. If a question would need extended reasoning, or weighs several independent factors, decompose it. Ask each factor separately and combine the answers with your own formula, so reprioritising later is a coefficient change in code rather than a prompt rewrite. Questions are evaluated in parallel and in isolation, so asking more barely changes response time and context does not rot across them.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| state | string | object | array | ✓ | The content to evaluate. A plain string for text, or structured data for chat logs, records, or the current state of your application. |
| model | string (optional) | ✓ | Optional. JevStation pins the model at the deployment level, so this field is ignored; TypeSafe's own API accepts the alias jev-latest. |
| questions | map<string, Question> | ✓ | A map of typed questions. You choose each key and the matching answer comes back under it; the key is a label only and is never sent to the model. Choice needs at least two options, Score at least two ordered levels. |
Every question object shares instructions (what to decide or rate) and adds its own criteria: an option map for Choice, an ordered level array for Score, and optional true/false descriptions for Noul.
Answers
One answer per question, keyed by the ids you supplied. Choice and Score answers carry a confidence value derived from their probability distribution; Noul answers return a probability only.
choiceThe highest-probability option, every option mapped to its probability, and a confidence value.
scoreA probability-weighted value across your levels (it can land between them), the legend mapping each level back to its description, and confidence.
noulThe probability that the answer is yes, from 0 (no) to 1 (yes). Noul answers carry no confidence.
Every response also reports token usage for the request as usage.input_tokens and usage.output_tokens.
Confidence
The shape of a probability distribution is what tells you how certain the model is: concentrated on one outcome is confident, spread across several is not. The confidence value collapses that shape into a single number from 0 to 1, so you can threshold on it without doing the maths yourself. A useful starting pattern is three bands:
Act automatically. The model has a clear read and you can proceed without human involvement.
Proceed with caution. Ask the user to confirm, flag for review, or gather more information first.
Do not act. Route to a human, ask for clarification, or fall back to a different system.
Where you draw the boundaries depends on the stakes: gate a destructive action at a higher threshold than a read-only one.
Full example
One call, three questions, one response — no SDK required.
const res = await fetch("https://your-deployment/api/v1/systemone", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.JEVSTATION_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
state: "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoicing, refunds",
technical: "Bugs, outages, integrations",
sales: "Pricing, upgrades, new accounts",
},
},
frustration: {
type: "score",
instructions: "How frustrated is the customer?",
criteria: ["Calm", "Frustrated", "Very angry"],
},
is_urgent: {
type: "noul",
instructions: "Does this convey urgency?",
},
},
}),
});
const { data } = await res.json();
data.answers.department.choice; // "billing"
data.answers.department.confidence; // 0.59
data.answers.frustration.score; // 1.04
data.answers.is_urgent.noul; // 0.99
data.credits.charged; // 1
data.credits.remaining; // 4,996curl -X POST https://your-deployment/api/v1/systemone \
-H "Authorization: Bearer $JEVSTATION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": "Payouts have been failing for 3 days.",
"questions": {
"is_urgent": { "type": "noul", "instructions": "Does this convey urgency?" }
}
}'Notes & limits
- An evaluation is a single round trip: Jev returns answers, not a token stream, so a run either produces a complete result or fails cleanly.
- Adding questions to a set costs extra input tokens but barely any extra latency, because they are evaluated in parallel against the same state.
- Each run consumes credits: one for a standard evaluation, three once the state passes 8,000 characters or the set passes five questions. The balance is checked before the model is called, and a failed or cancelled run is refunded.
- Evaluations are scoped to the signed-in user; one account can never read another's.
- Provider credentials are configured by the operator in the admin settings and never reach the browser.
- On 429 or 529, retry with exponential backoff. The client SDKs already do this for you.