LLM router comparison: LiteLLM vs OpenRouter vs RouteLLM vs Jev

“LLM router” covers two different jobs. Gateways such as LiteLLM, OpenRouter and Portkey decide which provider serves a model and handle retries and fallbacks. Routers such as RouteLLM and Jev decide which model a prompt needs in the first place. Most production setups want one of each.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Pick the cheapest model tier that can handle a request.

195 / 2,000

Edit freely — the model answers the scenario's questions about whatever you put here.

Model

At a glance

The tools below overlap in name more than in function. This table sorts them by the decision they make. Details are taken from each project’s documentation, which changes often, so check the linked docs before you commit to one.

LiteLLMOpenRouterRouteLLMPortkeyJev
KindGateway and SDKManaged gatewayLearned router (open source)Gateway and observabilityTyped classifier
Main decisionWhich deployment serves a model: load balance, retry, fall backWhich provider serves a model, plus an auto-routing optionStrong or cheap model, from a router trained on preference dataWhich provider, with fallbacks and controlsOne of your own tiers, plus scores and yes/no signals
Decides how hard a prompt isNot by defaultThrough its auto routerYes, that is its jobNot by defaultYes, you define the tiers
Calls the model for youYesYesYes, as a drop-in clientYesNo, it returns only the decision
Output you can thresholdNoNoA routing scoreNoA probability per tier, plus confidence
You run itSelf-host or hostedHostedSelf-hostHosted or self-hostHosted API

How each one works

LiteLLM. One interface to many LLM providers. Its Router supports weighted, latency-based, least-busy and cost-based strategies, with retries and fallbacks. It balances load across deployments of a model you have already chosen. It is the usual pick when you want to self-host the gateway.

OpenRouter. A hosted marketplace of models behind one API. It routes around provider outages, supports fallback chains of models, and offers an auto-routing option that picks a model for the prompt. It is the quickest route to many models without running anything.

RouteLLM. An open-source framework from LMSYS. Trained routers decide whether a prompt needs a stronger model or can go to a cheaper one. Its authors report large cost reductions at close to GPT-4 quality on their benchmarks; treat that as benchmark-specific, not as what your traffic will do.

Portkey. A gateway with routing, fallbacks, monitoring and access controls, aimed at teams that need governance around model use.

Jev. A classifier, not a gateway. You write the tiers, for example small, medium and frontier, each with a one-line description of the work it can handle. Jev returns a probability for each tier, a difficulty score and any yes/no signals you ask for, in one call. TypeSafe states 70 to 500 ms and $0.042 per million input tokens, with free output. On JevStation a three-question routing decision is 1 credit.

Choosing between them

  • You need failover, load balancing or one API across providers. Use a gateway: LiteLLM if you self-host, OpenRouter if you want it managed.
  • You want a router trained on preference data and have a strong model plus a cheap one. Look at RouteLLM.
  • You want to define the tiers and signals yourself and act on confidence. Use Jev, and read the probability.
  • You are unsure a router will beat a simple rule. Measure first. LiteLLM published a benchmark in which its rule-based complexity score reached an AUC of 0.524, close to random. Any router, Jev included, should beat a rule you could write in ten lines on your own labelled prompts.

Using a gateway and a classifier together

The two layers answer different questions, so they stack:

  1. Send the incoming prompt to Jev with a route Choice and any signals you care about, such as needs_code.
  2. Read the tier and its confidence. When confidence is low, go up a tier: a small waste of money costs less than a failed answer and a retry.
  3. Pass the request to LiteLLM, OpenRouter or the provider SDK with the model that tier maps to, and let the gateway handle retries and fallbacks.
  4. Log the tier, its probability and the model version Jev returned, so you can re-check thresholds when the version changes.

What the evidence says about Jev as a router

  • A DevelopersIO test measured $0.000025–$0.000027 per Jev call and a median of 0.643–0.674 s. It sent 40 calls across four difficulty tiers and got the expected tier each time.
  • That test used one prompt per tier and its author says it is not an accuracy benchmark. The medium tier came back with lower confidence than the extremes, which is where a router earns or loses its savings.
  • No public benchmark compares Jev, RouteLLM and the gateway auto-routers on the same prompts. Run your own.

You can try your own tiers and prompts in the LLM router tool or the playground, and see how Jev compares with a prompted LLM in Jev vs LLM classification.

Frequently asked questions

What is the difference between an LLM gateway and an LLM router?
A gateway gives you one API in front of many models or providers and handles load balancing, retries and fallbacks. A router decides which model a given prompt should go to, usually to save money on easy prompts. Some products do both, but the two jobs are separate, and a router can sit in front of any gateway.
Which LLM router is best?
It depends on the job. LiteLLM is a common choice for a self-hosted gateway and OpenRouter for a managed one. RouteLLM is the research option for learned strong-or-cheap routing. Jev suits the case where you want to define your own tiers and signals and read a probability for each. There is no single best tool, and no public benchmark compares them on the same traffic.
Can I use Jev with LiteLLM or OpenRouter?
Yes. Jev only returns a decision, such as a tier and its probability. Your code then sends the request to the model that tier maps to, through LiteLLM, OpenRouter or the provider directly.
How do I know a router is saving money?
Label a sample of your own prompts, compare the router’s tiers with your labels, and count how often the cheap tier was chosen and how often those answers needed a retry. A router that rarely picks the cheap tier, or whose cheap answers get retried, saves little.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.