Jev LLM router: send each prompt to the cheapest model that can handle it

Paste a user request. Jev rates how demanding it is, picks a model tier — small, medium or frontier — and says whether the answer needs code, all as probabilities your router can threshold. No LLM call sits in front of your LLM call.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Pick the cheapest model tier that can handle a request.

195 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

What this router checks

The example above is a single user request: a streaming parser for a 2GB CSV of sensor readings that has to find per-device gaps and stay under 500MB of memory. Jev answers three questions about it in one call:

QuestionTypeWhat you get back
difficultyScoreA position on Trivial → Easy → Moderate → Hard → Expert
routeChoiceA probability for small, medium and frontier, plus confidence
needs_codeNoulProbability that answering requires writing code

route is the decision; the other two are signals you log and can route on too. The request is deliberately a borderline one — the work is ordinary Python, but the memory limit and streaming requirement make it more than a lookup — so the interesting output is how the probability splits between medium and frontier, and how confident Jev is about the split.

Three questions on a short prompt cost 1 credit, and because Jev evaluates every question in a request in parallel, the extra two add almost no latency.

How to write routing questions that work

Describe tiers by the work, not the model name. Jev does not know what your "small" model can do. "A fast, cheap model is enough" is a start; "direct lookups, extraction and localized changes with explicit targets" — the wording from the LangChain router example — is better, because it names the kind of task.

{
  "type": "choice",
  "instructions": "Which model tier should handle this request?",
  "criteria": {
    "small": "Lookups, extraction, rewording and short edits with an explicit target",
    "medium": "Multi-step tasks and ordinary code with clear requirements",
    "frontier": "Novel reasoning, architecture, or code with hard constraints"
  }
}

Ask for the decision directly. A route Choice gives you one distribution to act on. Deriving the tier from difficulty alone means inventing cut-offs on a Score, which is harder to tune.

Add a Noul for each hard requirement. If code tasks must go to a code model, ask needs_code rather than hoping the tier description covers it. The same goes for "contains personal data" or "needs a long context".

Keep counting out of the question. TypeSafe documents numeric precision as a weakness. "Is the input over 10,000 tokens?" belongs in your code, not in a Jev question.

For more on phrasing, see Designing Questions Jev Can Answer.

From one check to a routing layer

  1. Route before the first completion. Call Jev with the incoming prompt, read route, and send the request to the model for that tier. TypeSafe’s docs list model routing as a use case in exactly these terms: classify intent and domain, estimate difficulty and risk, escalate what needs a stronger model.
  2. Use confidence as the second axis. TypeSafe’s confidence-gated routing pattern treats the choice as what and the confidence as whether to act. A sensible default for a router is to go up a tier, not down, when confidence is low: sending an easy prompt to a bigger model wastes a little money, while sending a hard one to a small model costs a failed answer and a retry.
  3. Log the distribution with the choice. When a run costs more than expected, the probabilities tell you whether the router was sure or guessing.
  4. Record the model version. Every response carries a model field such as jev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request.
  5. Call it from your code or your framework. The same evaluation runs through the System One API with an API key. In LangChain, ModelRouterMiddleware (experimental) picks a model from the latest user message and uses it for the whole run.

The wider pattern — routing, gating and escalation in one agent loop — is covered in Building a Harness with Jev.

Check it before you trust it

  • Label a sample of your own prompts. A Classmethod engineer ran 40 calls across four tiers and got the expected tier 10 out of 10 times for each, at a median of about 0.65 seconds and roughly $0.000025 per call. The write-up is explicit that it used one prompt per tier and is not an accuracy benchmark. Your traffic is the benchmark.
  • Watch the middle tier. In that same test the extreme tiers came back with confidence 1.0 while the medium tier sat between 0.57 and 0.67. Borderline prompts are where a router earns or loses its savings, so look at them first.
  • Compare against a dumb baseline. LiteLLM published a benchmark in which its rule-based complexity score reached an AUC of 0.524, close to random. If Jev’s tiers do not beat a rule you could write in ten lines, the router is not paying for itself.
  • Check the savings, not only the accuracy. Count how often small was chosen and how often those answers needed a retry. A router that is never wrong because it almost always picks frontier saves nothing.

Where Jev is the wrong router

  • The decision depends on numbers. Token counts, context length and price arithmetic are code, not classification.
  • You need the router to explain itself to a user or in an audit log. Jev returns probabilities, not a reason.
  • You want the gateway, not just the decision. Jev picks a tier; it does not proxy the request, retry or fail over. Pair it with the gateway or client library you already use, and let Jev decide only which model the request goes to.
  • Every request goes to the same model anyway. If one model is cheap enough for all your traffic, a router adds a call and saves nothing.

Try your own prompts and tiers in the playground, or see what a routing decision costs at volume on the pricing page.

Frequently asked questions

Why use Jev as a router instead of a small LLM?
A router runs in front of every request, so its own cost and latency add to all of them. Jev returns a typed choice in one round trip — TypeSafe states 70 to 500 ms and $0.042 per million input tokens, with free output — and it can only answer with one of the tiers you defined.
How accurate is Jev at routing?
An independent test by a Classmethod engineer sent 40 calls across four difficulty tiers and got the expected tier every time, but it used one prompt per tier and its author says it is not an accuracy benchmark. Measure routing on your own labelled prompts before you trust a threshold.
What does one routing decision cost on JevStation?
The example asks three questions about a short prompt, which is a standard evaluation: 1 credit. A standard evaluation covers up to 5 questions and 8,000 characters of state; above either limit it costs 3 credits.
Can I plug Jev into LangChain as a model router?
Yes. langchain-typesafe ships a ModelRouterMiddleware that picks a model from the latest user message and keeps it for the whole run. It is marked experimental, so its API may change.
Does Jev call the model it picks?
No. Jev only returns the decision and its probabilities. Your code, or a gateway, sends the prompt to the chosen model.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.