Intent classification: detect what a customer message is asking for

Paste a customer or chatbot message, list the intents you handle in plain English, and Jev returns a probability for each intent plus whether a person needs to act — no training data and no labelled utterances. Try it below on a banking message, without an account.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Detect what a customer message asks for, from intents you define.

126 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

Model

What this tool checks

The example is a message that fits two intents at once. The customer was charged twice — a duplicate charge — and one of the charges is still pending, which is its own intent in most banking flows. A classifier that returns a single label hides that. One evaluation asks two questions:

QuestionTypeWhat you get back
intentChoiceA probability for each of eight intents, including other, plus confidence
needs_humanNoulThe probability that a person, not an automated answer, has to act

The intents are written in the style of Banking77, a public dataset of banking customer messages, so the published results below are a reasonable reference point. None of them was trained: the line of text next to each intent is the whole definition.

Why probabilities matter for intent detection

A chatbot that routes on the top intent alone treats a 0.95 answer and a 0.45 answer the same way. With a probability for every intent you can do three different things:

  • Act automatically when one intent is far ahead — the confidence value is high.
  • Ask a clarifying question when two intents are close. "Do you want to cancel the pending charge or dispute the duplicate?" is better than guessing.
  • Hand off when other is likely or needs_human is high.

Because Jev can only answer with the options you defined, there is no invented intent name and no output to parse — the routing code reads a key and a number.

How it compares

The only independent measurement we know of is jev-benchmarks, a pre-registered pilot that ran jev-1.13.0 against the zero-shot model GLiNER on 100 held-out examples per dataset. On Banking77, with 72 intents in the sample:

ModelAccuracyMessages handled automatically at ≤5% errorMedian latency
Jev0.87086%246 ms
GLiNER0.61027%296 ms

GLiNER ran locally on an Apple M4 Max CPU and Jev was a hosted API called over the network, so the latency column compares two deployments, not two models. It is one 100-message pilot, not a leaderboard, and it did not include a prompted LLM or a classifier trained on Banking77, which would set the bar higher. On JevStation a Choice holds up to 12 options, so a 72-intent set like this one needs the two-stage layout described below.

How to define intents

Describe what the customer wants, not the words they use. "The customer was charged more than once for the same purchase" catches "double charged", "billed twice" and "two payments for one order" without listing phrases.

Separate intents that lead to different actions. If a duplicate charge and a pending payment go to the same flow, merge them. If they go to different flows, keep both and let the probabilities show when a message is about both.

Always include other. Its probability tells you how often real messages fall outside your flows — a direct signal of which intent to add next.

Go two levels deep for large sets. Ask a first Choice for the area — cards, transfers, account, other — then a second, area-specific Choice for the intent. Because one evaluation can ask up to 12 questions, you can send the area question and every area’s intent question together, then read only the intent question for the area that wins.

For more on phrasing, see Designing Questions Jev Can Answer.

From a test to a pipeline

  1. Pull a sample from your logs. A few hundred real messages, each labelled with the intent your team would assign, show where the definitions are ambiguous.
  2. Run them in one go. The batch page runs one question set over every row of a CSV, no code needed.
  3. Choose thresholds. Pick the confidence above which the bot acts alone, and the needs_human probability above which it hands off.
  4. Classify live messages through the API. The same evaluation runs through the System One API with an API key, at the same credits as the playground.

Where Jev is the wrong tool

  • You also need entity extraction. Jev classifies; it does not pull out the amount, merchant or date. Extract those with code or a generative model.
  • The answer depends on arithmetic. "Was I charged more than my limit?" is a comparison of numbers, a weakness TypeSafe documents. Compare in code.
  • You already have a large labelled set and stable intents. A classifier trained on your own data will usually beat a zero-shot model on that task.
  • Messages must not leave your infrastructure. Jev is a hosted model.

For routing whole support tickets by team and urgency rather than by intent, see the support ticket triage tool.

Frequently asked questions

What is intent classification?
Intent classification, or intent detection, maps a user message to the action it asks for — check a balance, cancel a payment, reset a password. Chatbots and helpdesks use it to decide which flow, answer or team handles the message.
Do I need training data for intent classification with Jev?
No. You define each intent in a line of plain English and Jev classifies against those definitions. You do need a few hundred labelled messages to measure accuracy and pick thresholds before relying on it, but not to make it work.
How accurate is Jev at intent classification?
In an independent, pre-registered pilot on 100 messages from Banking77, a banking intent dataset, Jev reached 0.870 accuracy against 0.610 for the zero-shot model GLiNER, and could handle 86% of messages automatically at an error rate of 5% or less. It is one small pilot; measure on your own messages.
How many intents can I define?
On JevStation one Choice question holds up to 12 intents. For a larger set, classify in two stages: a coarse category such as cards, transfers or account first, then the specific intent within that category.
Is this intent classification tool free?
You can try it free: the tool on this page runs without an account, 3 times a day, on an example you can edit up to 2,000 characters. To use your own intents, a free JevStation account includes 200 credits valid for 30 days, and one message costs 1 credit.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.