fast-jev-compaction: Compaction Without the Summary
A Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored, stale ones are dropped, and everything kept stays verbatim.

Every compaction scheme in a coding agent rests on the same bet: that a model can read a long transcript and write a shorter one that keeps whatever matters later. That bet loses. A file path, an exact error string, a constraint the user stated once in passing — these are cheap to keep and impossible to reconstruct once a summarizer has decided they were not worth the tokens.
fast-jev-compaction takes the opposite approach. Instead of rewriting the transcript, it asks a classifier which parts of it are still needed, deletes the ones that are not, and passes everything else through byte for byte.
Compaction is a selection problem, not a writing problem
Work through what a tool call is actually worth at compaction time. A Read from forty turns ago whose contents are still shaping the fix, a Grep that returned nothing useful and led nowhere, a failing test log that has since been superseded: each of these is a candidate for deletion, and each of them is a yes-or-no question.
The moment you phrase it that way, you have described a System One shaped task. Jev evaluates a state and a set of typed questions and returns probabilities, which means the decision is a number you can compare against a threshold instead of a paragraph you have to trust. LangChain's framing of why this matters is worth quoting: Jev "isn't a drop-in replacement for an LLM. It doesn't generate text, but it can handle classification tasks we often use LLMs for today."
The cost profile fits the job. By TypeSafe's published numbers Jev answers in one round trip, at 70 to 500 ms and $0.042 per million input tokens, so scoring every tool call in a session costs less than the single summarization call it replaces.
The whole design is three outcomes
For every tool call outside a pinned window, fast-jev-compaction asks Jev two yes/no questions and sorts the result against a threshold:
- keep — the call and its result stay exactly as they are;
- drop the result — the call stays, and the result shrinks to its first few hundred characters plus a note;
- drop the call — the call disappears, and its result goes with it.
That is the entire policy. Two decisions per call, three possible outcomes, one threshold. The code that implements it is small enough to read in one sitting, which is the point: a context management scheme you cannot audit is a context management scheme you cannot debug.
Step by step: what happens when it runs
-
Pair the calls. Every
tool_useblock is matched with itstool_resultbytool_use_id. A call with no result yet is not a candidate, because there is nothing to drop. Each pair gets a short id (t1,t2, …) used in the questions and in the logs. -
Pin what must not move. Calls in the first message and in the newest
preserveRecentMessagesmessages (6 by default) are pinned. Pinned calls are decided askeepbefore any question is built, so the opening instruction and the most recent work are never candidates. -
Build the state. This is the design decision worth pausing on. The transcript cannot be sent to Jev as-is: a real session is hundreds of thousands of characters, and Jev's practical request budget is around 32,000 tokens. So every tool result is replaced by a one-line note such as
ok, 4213 chars (omitted), while tool inputs and message texts are included. Nothing is summarized. The state carries acontextstring explaining the task, agoal(the last three user prompts by default), and the history oldest first. -
Fit the state into the budget. If it does not fit
maxStateTokens(25,000 by default), it is shrunk in ordered stages, each tried only because the previous one was not enough. More on this ladder below. -
Ask two questions per non-pinned call. The wording is fixed, and it is the most explicitly engineered part of the project:
Tool call t12 (Read) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does nextThe full output of tool call t12 (Read, 4213 chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not doThe second question is the interesting one, because "re-running the tool would not do" is exactly the criterion that separates a result worth keeping from one worth re-fetching.
-
Batch, concurrently. Calls are split into as many requests as needed so that the full state plus one batch of questions stays under
maxRequestTokens(30,000). The same complete state is resent with every request, the requests go out in parallel, and the answers are merged. -
Decide. Against
keepThreshold(0.5 by default):keepResultat or above it keeps everything; otherwisekeepCallat or above it keeps the call and truncates the result totruncateHeadChars(300) plus a note reading[fast-jev-compaction truncated 3913 chars of this tool result; re-run the tool if needed]; otherwise both go. If the state leaves no room for even one pair of questions, the run throws rather than guessing. -
Rebuild the transcript. Messages that lost content are rebuilt, messages that lost everything are removed, and untouched messages are returned as the same objects that came in. A result never survives its call. Nothing is reworded.
The part that makes it usable: shrinking the state, not the messages
The most instructive piece of engineering in this repository is not the scoring. It is what happens when the conversation is too large for Jev to see.
The library never silently loses a message. It downgrades its representation, in stages, and only as far as it must:
| Stage | What changes | Cost |
|---|---|---|
full | Nothing | Full detail |
inputs<=1000, then <=200, then <=60 | Tool inputs truncated to that many characters | Argument fidelity |
texts abridged | Long texts become head plus tail with [… N chars omitted …], oldest non-pinned first | Prose fidelity |
old messages collapsed | Old non-pinned messages become a single [… N chars omitted …] note | Prose presence |
old calls compacted | Old calls become one line each: t12 Read file_path=src/a.ts → ok 480ch | Input detail |
old messages left out | Old messages carrying no call are dropped from the state | Prose presence |
old calls merged | Runs of adjacent call-only entries fold into one entry each | Per-entry envelope |
If even the last stage does not fit, the call throws instead of proceeding on a partial view. The stage that ended up being necessary is reported in stats.stateStage, so you can tell from a log line whether your sessions are routinely running at full fidelity or perpetually scraping the bottom of the ladder.
Two details in that ladder are worth stealing for any classifier-backed system. Pinned messages are abridged last, so the newest work keeps its detail longest while old turns degrade first. And the state is always fitted to a ceiling comfortably below Jev's real request limit, which is why the library's token ceilings are 25,000 and 30,000 rather than 32,000.
Counting tokens without a tokenizer
A dependency-free library cannot ship a tokenizer, so it estimates: one token per six letters of a word, half a token per digit, nine tenths of a token per other symbol. The comment above the function explains why the simple approach fails:
Estimates tokens without a tokenizer: a word costs one token per six letters, a digit half a token, any other symbol nine tenths. Calibrated against the usage Jev reports for real transcripts, where it lands 2–18% above the true count; a plain characters-per-token ratio undercounts the JSON-heavy states by up to 40%.
Estimates that run slightly high are the right error to make. Running low means sending a request that Jev rejects, which is a failed compaction rather than a conservative one.
Using the library
The package is a typed ES module with no runtime dependencies and needs Node 18 or newer. The shape it works on is a subset of Claude Code's SessionMessage, so a transcript can be passed straight in:
import {
compactMessages,
reductionRatio,
type Message,
} from 'fast-jev-compaction';
const transcript: Message[] = [
{
role: 'user',
text: 'Fix the failing test. Never edit src/generated.',
toolUses: [],
},
{
role: 'assistant',
text: '',
toolUses: [
{
tool_use_id: 'toolu_1',
tool: 'Read',
input: { file_path: 'src/a.ts' },
},
],
},
{
role: 'user',
text: '',
toolUses: [],
toolResults: [{ tool_use_id: 'toolu_1', text: '…file…' }],
},
// …
];
const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.messages, result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
// Not worth it: keep the original transcript, or summarize instead.
}
result.decisions is the audit trail, one entry per paired call, carrying keepCall, keepResult, the action taken, and the reason (pinned, kept, result_dropped, call_dropped). result.stats reports message and character counts before and after, per-reason counts, the state size in estimated tokens, which fitting stage was needed, the request count, and elapsed milliseconds.
Below the top-level function, the pipeline is exported as parts: collectToolCalls, fitState, batchCalls, decideCall, applyDecisions, buildJevRequest, parseJevResponse, plus the JevClient and the one-method JevAsker interface. Implementing ask(state, questions) is all it takes to swap in your own transport, which is how this design gets reused by ports into other harnesses.
The options, and when to move them
| Option | Default | What it controls |
|---|---|---|
apiKey | TYPESAFE_API_KEY | TypeSafe key, read from the environment |
model | jev-latest | Jev model name |
baseUrl | https://api.typesafe.ai/v1/systemone | System One endpoint |
fetch | native fetch | Injectable fetch, used by the unit tests |
goal | last 3 user prompts | Task description included in the state |
keepThreshold | 0.5 | Minimum probability for a call or result to stay |
preserveRecentMessages | 6 | Newest messages never touched |
maxStateTokens | 25000 | Estimated ceiling for the state |
maxRequestTokens | 30000 | Estimated ceiling for state plus one batch of questions |
truncateHeadChars | 300 | Characters of a dropped result retained before its note |
Three of these are worth changing early. Raise preserveRecentMessages if your agent regularly reasons across the last dozen turns and you would rather spend the tokens on the recent window than on aggressive pruning. Raise keepThreshold if you would rather over-keep than re-run a tool, particularly on slow commands where re-reading is expensive. Set goal explicitly in a long automated run: the last three user prompts are a decent guess at what the session is doing, but a pipeline knows.
Installing it in Claude Code
The plugin replaces Claude Code's built-in compaction with the pruned transcript. Function hooks are early access and may change between releases, so the checked-in type reference is generated against Claude Code 2.1.274. Enable the surface wherever Claude Code runs — ~/.claude/settings.json is the durable place:
{
"env": {
"CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1",
"TYPESAFE_API_KEY": "<your key>"
}
}
Then add the repository as a plugin marketplace and install it:
claude plugin marketplace add tamaratran/fast-jev-compaction
claude plugin install fast-jev-compaction@fast-jev-compaction
The install prompts for the plugin options (key, thresholds, truncateHeadChars, and the rest). Leave them at their defaults to read TYPESAFE_API_KEY from the environment. Restart Claude Code or run /reload-plugins.
The plugin needs no publishing step to work from a checkout: the marketplace is just the repository's .claude-plugin/marketplace.json, so this runs the current working tree directly.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .
One caveat worth stating plainly: the README shows npm install fast-jev-compaction, but as of this writing the package is not on the npm registry. The library is consumed from the repository itself, and the plugin path above is the one that works out of the box. If you want this in OpenCode instead, the port opencode-fast-jev-compaction is published.
Plugin-only options
The plugin adds four settings on top of the library's:
| Option | Default | What it controls |
|---|---|---|
compactAtPercent | 60 | Context percentage at which turn.complete requests compaction |
minReductionRatio | 0.25 | Minimum reduction required to replace the history with the pruned version |
apiKey | — | Sensitive plugin option, alternative to TYPESAFE_API_KEY |
model | jev-latest | Jev model name |
minReductionRatio is the safety catch that makes the whole thing acceptable to run unattended. There is no point replacing a transcript with a version 4% shorter, and there is real risk in replacing it with one that threw away context for nothing. If the reduction does not clear the bar, the hook logs the fact and delegates to Claude Code's built-in compaction.
What you see when it runs
The outcome appears as a toast:
fast-jev-compaction: kept N/M messages, no summary (…)when the pruned history replaced the summary;fallback to built-in summary (…)when Jev failed, the key was missing, the history could not be fitted, or the reduction was too small.
The log line carries the reduction, per-reason counts, state size, and request count, and a per-call decisions: line records both probabilities. That last line is what turns threshold tuning from guesswork into reading a log: you can see exactly which calls landed near 0.5 and decide which way you would rather they fell.
Why this is better than a summary, and where it is not
The argument is not that Jev is a better writer than a summarizer. It never writes. The argument is that fidelity has a price and this design pays it in a currency you can audit.
A summary spends its budget on prose about the transcript. This spends it on decisions about the transcript, and leaves the transcript itself intact. Concretely, that means:
- Nothing is paraphrased, so nothing drifts. A constraint stated in turn 3 is still the literal string in the pruned history.
- Every deletion is attributable. Each removed call has a probability attached to the decision to remove it.
- Failures are visible and cheap. The library throws rather than truncating silently, and the plugin falls back to the built-in summary. A compaction that cannot be done well does not happen badly.
- It runs on a classifier's budget. One round trip per batch of questions, at a cost the Jev pricing model makes negligible next to a summarization call.
The honest limits are these. Only tool calls and results are candidates: text messages are never removed or shortened in the output, so a session that is long because of conversation rather than tool output will not shrink much this way. Token sizes are estimates, not tokenizer counts. A probability is not proof, and the assistant can always re-run a tool, which is precisely the recovery path the design assumes. Because the full state is resent with every request, a history near the state ceiling costs one request per handful of questions. And the project is three days old at the time of writing: it was created on 2026-09-17, and function hooks themselves are early access.
The privacy consideration deserves its own sentence rather than a footnote. The whole conversation goes to TypeSafe's API as state. Tool results are omitted from it, but tool inputs and message text are not, and if either can contain a secret your threat model has to account for that.
The reusable idea
Strip away the Claude Code specifics and what remains is a pattern worth copying for any agent harness:
- Send the classifier the whole conversation and ask it for decisions, not for a rewrite. A model that returns probabilities gives you a threshold you can tune and a log you can read. A model that returns prose gives you neither.
- Make patience the default and deletion the exception. Two questions per call, with the result kept unless the model is reasonably sure the assistant can re-fetch it.
- Shrink the representation, never the record. When the classifier's own context is the constraint, degrade detail in stages and report which stage you needed. Silent loss is the failure mode to design out.
- Fall back rather than fail. Replace the summary only when the reduction is real, and hand the session back to the host's own compaction when the numbers are not there.
If you want to see what the probabilities look like before wiring any of this up, the playground takes a state and a question set and returns the distributions directly. To go deeper on what makes a question set good, read Designing Questions Jev Can Answer; for where these decisions belong in an agent loop, read Building a Harness with Jev. The source is at tamaratran/fast-jev-compaction under MIT.