Jev sentiment analysis: review sentiment, rating and refund requests

Paste a product review. Jev returns whether it is positive, mixed or negative, which star rating the text implies, and the probability that the reviewer is asking for a refund — as numbers you can sort, filter and route on, not a summary someone has to read.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Read the sentiment, implied rating and refund intent of a review.

163 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

What this tool checks

The example above is the kind of review that breaks a positive/negative classifier. It praises the basic editing, complains that export crashed twice on a 40-page file, credits support for replying within a day, and ends by asking for a refund for this month. One evaluation asks three questions:

QuestionTypeWhat you get back
sentimentChoiceA probability for positive, mixed and negative, plus confidence
ratingScoreA position on 1 star → 5 stars, the distribution and confidence
refund_requestNoulThe probability that the reviewer asks for a refund, from 0 to 1

The three answers do different jobs. sentiment is for dashboards and trend lines, and giving it a mixed option keeps reviews like this one from being forced into a side. rating puts a number on reviews that arrive without stars — support emails, app-store replies, survey comments. refund_request is the one that needs a person: it turns a review into a task. All three are answered in parallel against the same state, so together they cost 1 credit.

How to write the questions

Define every sentiment option. The example gives each option a line of criteria — mixed is "Both praise and complaints". Without it, Jev has to guess whether a review with one small complaint is positive or mixed. Write the boundary your team actually uses.

Use the rating as a threshold, not a precise number. A Score returns a probability-weighted position on the scale. TypeSafe notes that Score levels are weak in numerical calibration, so do not read an expected value between 3 and 4 as "3.4 stars". Ask a threshold question instead: is this review at 2 stars or below?

One Noul per action you would take. A refund request, a cancellation threat, a bug report and a feature request each lead somewhere different. Give each its own Noul, and add criteria where the line is subtle:

{
  "type": "noul",
  "instructions": "Does the reviewer ask for a refund?",
  "criteria": {
    "true": "The reviewer asks for money back, a refund or a credit for a charge.",
    "false": "The reviewer complains about value or price but does not ask for money back."
  }
}

Do not ask the same thing twice in opposite forms. In TypeSafe’s own documentation, a question about a refund and its negation, asked separately, returned probabilities that summed to 1.19. Pick one wording and tune the threshold on it.

For more on phrasing, see Designing Questions Jev Can Answer.

From one review to a pipeline step

  1. Label a sample first. A few hundred of your own reviews, each marked with the sentiment, rating and actions your team would assign, are enough to see where Jev’s answers match.
  2. Run the backlog in one go. The batch page runs one question set over every row of a CSV, so you can score an export of past reviews and compare the results with your labels without writing code.
  3. Route on the Nouls. Send reviews with a high refund_request probability to support, and use sentiment and rating for reporting. Up to five questions on a review under 8,000 characters still cost 1 credit, so adding a bug-report or cancellation question costs nothing extra.
  4. Record the model version. Every response carries a model field such as jev-1.13.0. Store it with each result and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request.
  5. Classify new reviews as they arrive. The same evaluation runs through the System One API with an API key, at the same credits as the playground.

Check it before you trust it

  • Be honest about the evidence. The only independent benchmark we know of is a pre-registered pilot of 100 examples per dataset. Jev reached 0.910 accuracy on AG News (four news topics), but only 0.480 on DAIR Emotion, a six-label emotion dataset, where it could not handle any share of inputs automatically at a 5% error rate and was worse calibrated than the zero-shot model it was compared with. Three-way sentiment is a different task from six-way emotion, and nobody has published a Jev result for it that we can cite. Test on your own reviews.
  • If you want emotions, test them hardest. Labels such as joy, anger, fear and sadness are close to the case that went badly. Measure them separately before building on them.
  • Check each language. TypeSafe says accuracy is best in English. Other languages are handled but not equally well.
  • Watch the mixed band. If most reviews land there, either your reviews really are mixed or the criteria are too vague — read a sample to find out which.

Where Jev is the wrong tool

  • You need a summary of what customers say. Jev returns typed answers, not text. Pair it with a generative model for summaries.
  • Fine-grained emotion is the product. That was the weak case in the benchmark above. A model trained on your labels may do better; Jev vs open models covers the options.
  • The rule is about numbers. "Has this customer left three negative reviews this month?" is counting, a weakness TypeSafe documents. Count in code.

If you are choosing between Jev, a general LLM and a classifier you train yourself for this job, Jev Alternatives sets out when each one is the better tool.

Frequently asked questions

How accurate is Jev for sentiment analysis?
There is no public benchmark of Jev on product-review sentiment that we can point to. In an independent pilot, Jev did well on news topics and banking intents but poorly on a six-label emotion dataset, so measure it on a few hundred of your own labelled reviews before relying on it.
What does classifying one review cost?
On JevStation, a review under 8,000 characters with up to five questions is a standard evaluation and costs 1 credit, so the three questions here cost 1 credit per review. Longer text or more questions cost 3 credits.
Can I classify thousands of reviews at once?
Yes. The batch page runs one question set over every row of a CSV, or you can call the System One API with an API key from your own pipeline.
Does it work on reviews that are not in English?
TypeSafe says English is where Jev is most accurate and that other languages, including CJK scripts, are handled but not equally well. Label a sample in each language you receive and check them separately.
Can Jev summarise what customers complain about?
No. Jev does not generate text. It can tell you whether each review mentions a crash, pricing or support, if you ask a question for each, but writing the summary needs a generative model.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.