Jev context compaction: decide which turns an agent must keep

Paste an excerpt of an agent transcript. Jev says whether it contains a decision or constraint the agent must remember, and rates how important it is to keep — as probabilities your compaction code can threshold, instead of a summary that rewrites what was said.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Decide which parts of an agent transcript are worth keeping.

332 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

What this check scores

The example above is four turns from a coding agent’s session. The agent reports a migration where 3 of 40 tables failed on a foreign key; the user mentions a spilled coffee; the agent offers to drop the constraint; the user says never to drop constraints in production and to backfill the missing customers first. Jev answers two questions about the whole excerpt:

QuestionTypeWhat you get back
keep_turnNoulProbability the excerpt holds a decision or constraint to remember
importanceScoreA position on Drop → Low → Medium → High → Must keep, plus confidence

The excerpt mixes the most important line in the session with the least important one. A summary might keep "backfill first" and lose "never drop constraints in prod"; Jev does not rewrite either, it only says how strongly the excerpt as a whole should stay. That is also the lesson of the example: if you want the coffee line gone and the constraint kept, score the turns one at a time.

How to write compaction questions that work

Ask about the future, not the past. "Is this turn relevant?" is true of almost everything. "Does the agent still need this to decide what to do next?" separates live context from history.

Give the reason a deletion is safe. The plugin behind fast-jev-compaction asks whether a tool’s full output should stay verbatim because "re-running the tool would not do". That phrase is the whole criterion: a result the agent can fetch again is cheap to drop, a user constraint is not.

Put the goal in the state. Whether a turn matters depends on the task. Include the last few user prompts, or the pipeline’s own task description, alongside the history.

Use a Score when you shorten rather than delete. A Noul gives keep or drop. A five-level Score lets you keep High and Must keep verbatim, truncate Medium, and drop the rest.

For more on phrasing, see Designing Questions Jev Can Answer.

From one check to a compaction step

The fast-jev-compaction plugin for Claude Code is a worked example of turning this check into a pipeline step:

  1. Pin. Calls in the first message and in the newest 6 messages are kept without asking.
  2. Build a small state. Every tool result is replaced by a one-line note such as its size, while tool inputs and message text stay in. Nothing is summarized.
  3. Ask two questions per remaining tool call — should the call stay, and should its full output stay — batched so each request stays under its token ceiling, sent in parallel.
  4. Decide at a threshold of 0.5 by default: keep both, keep the call with a 300-character head of its output plus a note to re-run the tool, or drop both.
  5. Replace the history only if it is worth it. Below a 25% reduction, or if Jev fails, the plugin hands back to Claude Code’s built-in summary.

The full design, including how it shrinks the state when a session is too large for Jev to see, is in the deep dive: fast-jev-compaction: Compaction Without the Summary. The same questions work from any harness through the System One API.

Check it before you trust it

  • Read the decisions, not just the savings. Log each probability next to the action. Items landing near your threshold are where tuning happens.
  • Replay real sessions. Take transcripts where the agent later needed something from early on, compact them, and check it survived.
  • Record the model version. Every response carries a model field such as jev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request.
  • Keep the state focused. TypeSafe lists a large state full of irrelevant detail as a known weakness. Filter in code first, and score in smaller pieces rather than one long excerpt.

Where Jev is the wrong tool

  • The session is long because of conversation, not tool output. Dropping whole tool calls is easy to make safe; cutting prose is not, which is why the plugin never shortens text messages.
  • You need a summary for a human reader. Jev does not write one.
  • The keep decision depends on dates or counts, such as "anything older than yesterday". TypeSafe documents date comparison and arithmetic as weaknesses; do those in code.
  • The context cannot leave your infrastructure. The state goes to the API.

Try it on your own transcripts in the playground, or see what scoring every turn costs on the pricing page.

Frequently asked questions

How is this different from summarizing the context?
A summary rewrites the transcript, and a constraint stated once in passing can be lost or paraphrased. With Jev you only decide what to delete; everything you keep stays word for word, and every deletion has a probability attached.
Can Jev read a whole long session at once?
Not always. TypeSafe gives a practical request budget of about 32,000 tokens, and JevStation caps the state at 24,000 characters. Long sessions need a shortened representation, such as replacing tool outputs with one-line notes, or several requests.
What does scoring one excerpt cost?
The example asks two questions about a short excerpt, a standard evaluation: 1 credit. State over 8,000 characters or more than 5 questions makes it a large evaluation at 3 credits.
Is there a ready-made plugin?
fast-jev-compaction is an open-source Claude Code plugin that asks Jev two yes/no questions per tool call and drops what is no longer needed. Function hooks, which it relies on, are early access in Claude Code.
Does my transcript leave my machine?
Yes. Whatever you put in the state is sent to the API as input. If tool inputs or messages can contain secrets, strip them before the call.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.