Back to blog

Building a Harness with Jev

An agent loop is only as fast as its slowest decision. Where Jev, the System One model from TypeSafe, fits: model routing, action gating, escalation.

Updated JevStationJevStation
Building a Harness with Jev

An agent is a loop: the model decides what to do, a tool runs, something checks the result, and the loop goes round again. The loop's speed is set by its slowest decision — and today almost every decision in that loop is a full call to a chat model.

Tool calling and structured outputs fixed the interface problem. Models can request tools and return data your code can parse. They did not fix the cost problem, because a decision that arrives as generated text is still a token-by-token completion. When LangChain wrote up building a harness with Jev, they made the same point from the framework side: the loop stays slow and expensive as long as every decision needs another model call.

Jev takes the other half of that work. It is a System One model from TypeSafe: send a state and a set of typed questions, get typed answers with probabilities back. No prose, no parsing, no stream to reassemble.

This post is the operator's version of the same story — where the Jev calls go in a harness, and what to do with the numbers that come back.

The decisions hiding inside the loop

A loop that looks like "think, act, observe" is really a sequence of small judgements:

  • Is the task finished, or does the agent need another turn?
  • Is this request simple enough for a cheap model?
  • Is this tool call safe to execute without asking a human?
  • Which of these five tickets is actually urgent?

Each one is a place where a chat model would generate an answer your code then has to interpret. Jev answers all four directly, which means you pay classification prices instead of generation prices. TypeSafe reports up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks — a vendor claim worth measuring against your own traffic, but the order of magnitude is what makes the pattern practical.

What comes back

A System One model is trained with reinforcement learning for calibrated decisions (RLCD), and calibration is the point: the probability it returns is meant to be used as a probability, not read as a vibe. You send one state — text, a JSON record, or a message list — plus any number of typed questions about it.

{
  "model": "jev-latest",
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
  "questions": {
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}
{
  "is_urgent": { "type": "noul", "noul": 0.999 }
}

Three question types cover most of what a harness needs to decide:

  • Choice picks one option from a list you define, and returns the probability of every option plus a confidence value.
  • Score rates the state against ordered levels, and returns a probability-weighted score, the distribution, and confidence.
  • Noul answers a yes/no question with a single probability between 0 and 1. It carries no confidence value, because the probability is the answer.

Two properties make this cheap to run inside a control loop. Questions are evaluated in parallel against the same state, so a set of six costs barely more wall-clock time than a set of one — you pay input tokens for the extra questions and almost nothing in latency. And an evaluation is a single round trip: you get complete answers or a clean failure, never a half-parsed stream.

Pattern 1: route before you spend

The cheapest way to speed up a loop is to stop sending easy work to an expensive model. Model routing is a Jev question — "what kind of task is this?" — asked before the first completion.

from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
    ModelChoice,
    ModelRouterMiddleware,
)

router = ModelRouterMiddleware(
    choices={
        "fast": ModelChoice(
            model="openai:luna",
            criteria="Direct lookups, extraction, and localized changes.",
        ),
        "powerful": ModelChoice(
            model="openai:sol",
            criteria="Architecture and high-stakes decisions.",
        ),
    },
    instructions="Choose the least costly model that can complete the task.",
)

agent = create_agent("openai:gpt-5.6-luna", middleware=[router])

The router classifies the latest user message once, and the run continues on the model it picked. Log the probability distribution next to the choice and "why did this run cost so much" becomes a query instead of an investigation.

Pattern 2: gate the action before the tool runs

The harder problem is trust. An agent that can run bash can be talked into running the wrong thing, either by a confused user or by instructions hidden in a file it just read. Coding harnesses such as Claude Code, Codex, and Cursor ship classifiers that judge an action before it happens; until recently that filter lived in closed-source harness code.

That check is not free — it is an extra round trip in front of every gated call, which is why those harnesses scope and sample it. The interesting thing about Jev is that it makes the check cheap enough to run on every call instead.

AutoModeMiddleware is the same pattern applied to any agent: Jev checks tool calls for risky decisions and blocks them before the tool executes.

from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import AutoModeMiddleware

guardrail = AutoModeMiddleware(tools=["bash"])

agent = create_agent("openai:gpt-5.6-luna", middleware=[guardrail])

Placement is the whole point. A guardrail that runs after the tool has executed is a log line. A guardrail that runs before it is a control. Because Jev answers in a single fast round trip, you can afford to check every call instead of sampling a few.

Two details matter before you lean on it. The middleware refuses risky calls rather than asking for approval, so if you would rather a person decide, pair it with human-in-the-loop middleware. And it only classifies the tools you name, which makes the guardrail exactly as complete as that list.

Pattern 3: decide what to do when confidence is low

Probabilities only matter if they change behaviour. A workable default is three bands, with the boundaries set by what the action costs when it is wrong:

  • High confidence — act automatically. The distribution is concentrated and the run can proceed without a human.
  • Medium confidence — proceed with a caveat: confirm with the user, collect one more piece of context, or queue the action for review.
  • Low confidence — do not act. Route to a human, ask a clarifying question, or fall back to a deterministic path.
const CONFIDENCE = { auto: 0.9, review: 0.6 };

function gate(answer: { choice: string; confidence: number }) {
  if (answer.confidence >= CONFIDENCE.auto) return { action: answer.choice };
  if (answer.confidence >= CONFIDENCE.review) {
    return { action: 'review', proposed: answer.choice };
  }
  return { action: 'escalate' };
}

Gate a destructive tool at a higher threshold than a read-only one. The numbers are policy, not physics, and they should be the easiest thing in your system to change.

One operational note that is easy to miss: an alias such as jev-latest moves when a new release ships, so the answers your thresholds were tuned against can shift without any change on your side. The response reports the versioned id that answered (jev-1.13.0 today), so log it — and once your bands are tuned, pin that id instead of the alias, and let a model upgrade be a decision rather than a surprise.

Running it without a framework

None of the above requires LangChain. A harness that already speaks HTTP can call Jev directly — which is exactly what the JevStation playground does. Each evaluation is one System One call whose state is the conversation so far, and the answers render as distributions instead of sentences.

The workflow is the one you would build anyway:

  1. Put the state in — a support ticket, a review, a spec, or a JSON record.
  2. Write the question set — mix Choice, Score, and Noul in the same call. The ids are yours to choose, and the answers come back under them.
  3. Read the numbers — every answer arrives with its distribution, plus a confidence value for Choice and Score.
  4. Move the call into your code — POST the same state and questions to the TypeSafe API with your own key, or keep tuning the questions in the hosted playground. The request shape is documented in the Jev docs.

In JevStation each evaluation costs one credit, checked before the model is called and refunded if a run fails or is cancelled, so a runaway loop cannot quietly empty an account.

What Jev is not

Jev does not replace the model that writes your agent's final answer. It cannot draft the email, explain the bug, or summarize the document — it does not generate text at all. The useful shape is a split: an LLM for open-ended reasoning and generation, Jev for the decisions in between, where the result is a branch rather than a paragraph.

Two honest caveats. First, the question set is the ceiling — a vague question returns a flat distribution, which is the model telling you it cannot tell the difference. Second, confidence is not correctness: a confidently wrong classifier is still wrong, so start with thresholds that escalate more often than you think you need, and tighten them once you have logs.

Calibration is also a property of a group of answers rather than a promise about any single one, and independent pilots have already found it uneven across datasets. Treat confidence as a strong prior, measure it against your own traffic, and keep a human path for the cases where being wrong is expensive. There is more on writing the questions themselves in Designing Questions Jev Can Answer.

Takeaways

  • Treat every branch in the loop as a decision rather than a generation: routing, gating, triage, and completion checks are all classification questions.
  • Put the cheap classifier before the expensive thing it protects — before the model call, before the tool executes.
  • Let confidence drive behaviour through explicit bands, and keep the thresholds in configuration, not code.

If you want to see the answers before you wire them into anything, paste a real ticket, log line, or spec into the playground and ask two or three questions about it. LangChain's original write-up is worth reading alongside this one for the framework view.