Jev context compaction: decide which turns an agent must keep
Paste an excerpt of an agent transcript. Jev says whether it contains a decision or constraint the agent must remember, and rates how important it is to keep — as probabilities your compaction code can threshold, instead of a summary that rewrites what was said.
自分のテキストで試す
例を編集して実行してください。アカウントは不要で、1 日 3 回まで無料で実行できます。
シナリオ
エージェントのトランスクリプトのうち、残す価値のある部分を判断します。
自由に編集できます。ここに入力した内容について、Jev がシナリオの質問に回答します。
What this check scores
The example above is four turns from a coding agent’s session. The agent reports a migration where 3 of 40 tables failed on a foreign key; the user mentions a spilled coffee; the agent offers to drop the constraint; the user says never to drop constraints in production and to backfill the missing customers first. Jev answers two questions about the whole excerpt:
| Question | Type | What you get back |
|---|---|---|
keep_turn | Noul | Probability the excerpt holds a decision or constraint to remember |
importance | Score | A position on Drop → Low → Medium → High → Must keep, plus confidence |
The excerpt mixes the most important line in the session with the least important one. A summary might keep "backfill first" and lose "never drop constraints in prod"; Jev does not rewrite either, it only says how strongly the excerpt as a whole should stay. That is also the lesson of the example: if you want the coffee line gone and the constraint kept, score the turns one at a time.
How to write compaction questions that work
Ask about the future, not the past. "Is this turn relevant?" is true of almost everything. "Does the agent still need this to decide what to do next?" separates live context from history.
Give the reason a deletion is safe. The plugin behind fast-jev-compaction asks whether a tool’s full output should stay verbatim because "re-running the tool would not do". That phrase is the whole criterion: a result the agent can fetch again is cheap to drop, a user constraint is not.
Put the goal in the state. Whether a turn matters depends on the task. Include the last few user prompts, or the pipeline’s own task description, alongside the history.
Use a Score when you shorten rather than delete. A Noul gives keep or drop. A five-level Score lets you keep High and Must keep verbatim, truncate Medium, and drop the rest.
For more on phrasing, see Designing Questions Jev Can Answer.
From one check to a compaction step
The fast-jev-compaction plugin for Claude Code is a worked example of turning this check into a pipeline step:
- Pin. Calls in the first message and in the newest 6 messages are kept without asking.
- Build a small state. Every tool result is replaced by a one-line note such as its size, while tool inputs and message text stay in. Nothing is summarized.
- Ask two questions per remaining tool call — should the call stay, and should its full output stay — batched so each request stays under its token ceiling, sent in parallel.
- Decide at a threshold of 0.5 by default: keep both, keep the call with a 300-character head of its output plus a note to re-run the tool, or drop both.
- Replace the history only if it is worth it. Below a 25% reduction, or if Jev fails, the plugin hands back to Claude Code’s built-in summary.
The full design, including how it shrinks the state when a session is too large for Jev to see, is in the deep dive: fast-jev-compaction: Compaction Without the Summary. The same questions work from any harness through the System One API.
Check it before you trust it
- Read the decisions, not just the savings. Log each probability next to the action. Items landing near your threshold are where tuning happens.
- Replay real sessions. Take transcripts where the agent later needed something from early on, compact them, and check it survived.
- Record the model version. Every response carries a
modelfield such asjev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request. - Keep the state focused. TypeSafe lists a large state full of irrelevant detail as a known weakness. Filter in code first, and score in smaller pieces rather than one long excerpt.
Where Jev is the wrong tool
- The session is long because of conversation, not tool output. Dropping whole tool calls is easy to make safe; cutting prose is not, which is why the plugin never shortens text messages.
- You need a summary for a human reader. Jev does not write one.
- The keep decision depends on dates or counts, such as "anything older than yesterday". TypeSafe documents date comparison and arithmetic as weaknesses; do those in code.
- The context cannot leave your infrastructure. The state goes to the API.
Try it on your own transcripts in the playground, or see what scoring every turn costs on the pricing page.
よくある質問
- How is this different from summarizing the context?
- A summary rewrites the transcript, and a constraint stated once in passing can be lost or paraphrased. With Jev you only decide what to delete; everything you keep stays word for word, and every deletion has a probability attached.
- Can Jev read a whole long session at once?
- Not always. TypeSafe gives a practical request budget of about 32,000 tokens, and JevStation caps the state at 24,000 characters. Long sessions need a shortened representation, such as replacing tool outputs with one-line notes, or several requests.
- What does scoring one excerpt cost?
- The example asks two questions about a short excerpt, a standard evaluation: 1 credit. State over 8,000 characters or more than 5 questions makes it a large evaluation at 3 credits.
- Is there a ready-made plugin?
- fast-jev-compaction is an open-source Claude Code plugin that asks Jev two yes/no questions per tool call and drops what is no longer needed. Function hooks, which it relies on, are early access in Claude Code.
- Does my transcript leave my machine?
- Yes. Whatever you put in the state is sent to the API as input. If tool inputs or messages can contain secrets, strip them before the call.
関連
パイプラインに組み込む
新規登録で 200 クレジットを無料進呈。独自の質問セットを保存し、同じ評価を API から呼び出せます。