Zero-shot classification: label text with categories you define
Write your labels as plain-English definitions, paste a text, and Jev returns a probability for every label plus a confidence value — no training data, no fine-tuning and no prompt to parse. Try it below on a news snippet, without an account.
Try it on your own text
Edit the example and run it. No account needed — 3 free runs a day.
Scenario
Sort a text into topics you define, with no training data.
Edit freely — Jev answers the scenario's questions about whatever you put here.
What this tool checks
The example looks like a boundary case: a company’s share price and revenue forecast, which is business, but the company makes processors, which could pull towards science and technology. With business defined as companies, markets and earnings, Jev puts it firmly in business. Rewrite the text to be about the processor itself and watch how the distribution moves. One evaluation asks three questions:
| Question | Type | What you get back |
|---|---|---|
topic | Choice | A probability for world, sports, business and sci_tech, plus confidence |
opinion | Noul | The probability that the text is commentary rather than a news report, from 0 to 1 |
negative_news | Noul | The probability that it reports a setback for the organisation it covers |
The four topics are the label set of AG News, a standard topic-classification benchmark, so you can set what you see here against the published numbers below. Nothing was trained for them: the only thing that defines each label is the line of text next to it. Change those lines and the classifier changes with them.
How zero-shot classification works here
Most zero-shot classifiers turn the task into something the model already knows how to do. The classic approach, used by Hugging Face’s zero-shot-classification pipeline, turns each label into a hypothesis — by default This example is {}. — and scores it with a natural-language-inference model. Prompting a general LLM to "answer with one of these labels" is the other common route.
Jev takes the labels directly. A Choice question is a list of options with criteria, and the answer comes back as a probability for every option, not as text you have to match against your label names. Two consequences matter in practice:
- It cannot answer outside your labels. There is no "Business/Tech" invented on the fly and no parsing step to fail.
- You get a distribution, not just a winner. A text at 0.55 business and 0.40 sci_tech is a different decision from one at 0.97 business. The confidence value condenses that shape into one number you can threshold.
How it compares with other zero-shot classifiers
The only independent measurement we know of is jev-benchmarks, a pre-registered pilot that ran jev-1.13.0 against GLiNER (fastino/gliner2.5-multi-v1) on 100 held-out examples from each of three datasets:
| Dataset | Labels | Jev accuracy | GLiNER accuracy | Jev coverage at ≤5% error | GLiNER coverage |
|---|---|---|---|---|---|
| AG News | 4 | 0.910 | 0.700 | 0.830 | 0.240 |
| Banking77 | 72 | 0.870 | 0.610 | 0.860 | 0.270 |
| DAIR Emotion | 6 | 0.480 | 0.440 | 0.000 | 0.020 |
Coverage is the share of texts you could act on automatically while keeping errors at or below 5% — closer to how a classifier is used than raw accuracy. Jev is clearly ahead on news topics and banking intents. On emotions neither model is usable, and Jev was the worse calibrated of the two. It is a 300-example pilot, not a leaderboard, and it did not include NLI models such as bart-large-mnli or a prompted LLM. Jev Alternatives covers the trade-offs in more depth.
How to write the labels
Define every label, not just name it. "Business" means different things to an editor and an accountant. A line of criteria — "Companies, markets, earnings and the economy" — is where your definition of the task lives.
Make the boundary explicit. If your team would file a chipmaker’s earnings under business, say so in the criteria: "Company results belong here even when the company is a technology firm."
Add an escape option. Real data contains texts that fit no label. An other option keeps them from being forced into the nearest one, and its probability tells you how often your label set misses.
Use Nouls for tags that overlap. If a text can be about pricing and about bugs at the same time, that is multi-label classification. Ask one Noul per tag instead of one Choice.
For more on phrasing, see Designing Questions Jev Can Answer.
From a test to a pipeline
- Label a sample first. A hundred or so texts your team has already classified are enough to see where Jev agrees and where your label definitions are ambiguous.
- Classify a backlog without code. The batch page runs one question set over every row of a CSV.
- Pick a threshold from the sample. Act automatically on answers above it and send the rest to a person — that split is what the coverage column above measures.
- Classify new texts through the API. The same evaluation runs through the System One API with an API key, at the same credits as the playground.
Where Jev is the wrong tool
- Fine-grained emotion. That was the failure case in the benchmark. Test it hardest, or train a model on your own labels.
- The text must stay on your machines. Jev is a hosted model. GLiNER or an NLI model can run locally.
- You have a large, stable label set and plenty of labelled data. A classifier trained on your data will usually beat any zero-shot model on its own task, at a lower cost per call.
- You need the reasoning written down. Jev returns typed answers, not explanations.
Frequently asked questions
- What is zero-shot classification?
- Zero-shot classification assigns a text to labels the model was never trained on for that task. Instead of collecting labelled examples and training a classifier, you describe each label and the model judges which one fits. It trades some accuracy on narrow, stable tasks for being usable the moment you write the labels down.
- Is this zero-shot classifier free?
- You can try it free: the tool on this page runs without an account, 3 times a day, on an example you can edit up to 2,000 characters. To classify your own texts into your own labels, a free JevStation account includes 200 credits valid for 30 days; one text with up to five questions costs 1 credit.
- How does Jev compare with GLiNER for zero-shot classification?
- In an independent, pre-registered pilot of 100 examples per dataset, Jev scored 0.910 accuracy on AG News against 0.700 for GLiNER, and 0.870 against 0.610 on Banking77. On the six-label DAIR Emotion set neither was usable. GLiNER runs locally and answered faster on small label sets.
- How many labels can I use?
- On JevStation one Choice question holds up to 12 options, and one evaluation can ask up to 12 questions. For a larger label set, classify in two steps: a coarse category first, then the label within it.
- Can a text have more than one label?
- Yes, with a different question type. A Choice picks exactly one option. For multi-label classification, ask one Noul per label — "Does this text discuss pricing?" — and each returns its own probability from 0 to 1.
Related
- Intent classificationDetect the intent of a customer or chatbot message from intents you define, with a probability for each.
- Review sentimentClassify a product review as positive, mixed or negative, estimate its star rating and flag refund requests.
- Jev vs open modelsopen-jev, jev-on-a-laptop, GLiNER and fine-tuned classifiers against hosted Jev: what each is, and what has been measured.
Make it part of your pipeline
Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.