Context compaction made our agent blame the user
Our agent runtime compacts its own context, and when it has to cut a running turn short it calls the Claude Agent SDK's interrupt(). The SDK writes that interrupt into the transcript as the user's decision. After one of those cuts, our agent told a person they had stopped it while it was listing and downloading files. Nobody had.
The bug wasn't in the model. It was in a sentence the SDK wrote into the transcript on our behalf, and in a part of our runtime that believed it.
Why our runtime compacts context itself
Native compaction in the Claude Agent SDK works by summarizing. Anthropic's agent loop documentation puts it plainly: "When the context window approaches its limit, the SDK automatically compacts the conversation: it summarizes older history to free space, keeping your most recent exchanges and key decisions intact." For a coding session that's usually fine. For an operations agent that has been working with someone all day, the summaries in our sessions still lost the things that matter most: the exact file and sheet IDs it was editing, a rule like "formatting only, don't touch the values", and what it had already delivered.
So since June 2026 (fleet-wide since July) our runtime owns compaction. It polls getContextUsage() on every assistant message. At 70%, the default for our Claude agents, it sets a flag and lets the current turn finish. Then a separate model call extracts a structured handoff ledger from the transcript (open task, IDs, rules, what's done, what's waiting on whom), and a fresh session starts from that ledger instead of from a summary.
The "let the turn finish" part is a lesson I paid for. In a June canary the runtime interrupted a turn mid-tool, crashed before the reseed could run, and the user watched a typing indicator for 17 minutes with no reply. Flag-and-wait removed that whole class of failure.
There is one exception, plus a preflight check that only applies to our non-Anthropic endpoints. If a turn keeps running and the gauge reaches 78%, a backstop calls interrupt() so the session can't run past the point where native compaction would take over.
What happened
Between October 3 and October 8 I traced five backstop interrupts on one client's operations agent. Each landed 42 to 64 milliseconds after the agent had called a file-storage tool, before the tool ran.
The SDK then wrote two things into the transcript. The tool result for that call said:
The user doesn't want to proceed with this tool use. The tool use was rejected (eg. if it was a file edit, the new_string was NOT written to the file). STOP what you are doing and wait for the user to tell you how to proceed.
And the next record was a marker beginning [Request interrupted by user.
Our ledger extractor did exactly what it was built to do: it read the transcript and recorded the facts. The file step went into the ledger as rejected, with the next move waiting on the user. Each time, a fresh session started from that ledger two to five minutes later. In the case that got flagged on October 8, the agent did the polite thing: it told the person it was working with that they had stopped it.
| What the transcript said | What actually happened |
|---|---|
| The user doesn't want to proceed with this tool use | Our runtime cut the turn to compact context |
| Request interrupted by user | No person touched anything |
| Step rejected, waiting on the user | The file call never ran and was still needed |
Since October 1, ten sessions across two agents carried that rejection text. I traced five events to the backstop. The same text also follows a person refusing a permission prompt, and I haven't traced the other sessions, so some of the ten may not be this bug.
One more detail from the same events. At the moment the backstop fired, the API's own usage for the last request was between 67% and 77% of the 200K window the gauge measured against, while the gauge had crossed 78%. The TypeScript reference describes the /context total, which getContextUsage() returns by default, as "Claude Code's estimate of the tokens in use". In these five events it read above the last API call. Elsewhere, on another backend, it read 70% while the request going out was near the limit, so treat it as a gauge with its own error, not as the bill.
Why the transcript can't tell you who interrupted
interrupt() takes no arguments, so it can't carry a reason, and no documented field in the transcript or the result says who issued it. The TypeScript reference says it "Interrupts the query" and, on newer CLIs, returns a receipt of queued messages. Nothing in it records why. The Python reference confirms that "the tool result carries the interruption message", and the result's terminal_reason is aborted_tools for an interrupt() and for a permission callback that denies with interrupt=True alike.
In our runtime a person can only stop a turn through our own code, so the only code that knows why interrupt() was called is ours. The transcript does carry an undocumented toolDenialKind: "interrupted", and the SDK changelog describes a non_execution_kind field added so consumers can "classify denied, interrupted, or cancelled tool calls without string-matching result prose". Neither is in the public TypeScript types of 0.3.282, the version I run in production, or of 0.3.296, the latest. The changelog entries from 0.3.283 to 0.3.296 say nothing about interrupt attribution.
If you call interrupt() from your own code, assume the transcript will say a person did it.
The fix
The fix has three parts, and none of them changes the SDK.
First, the runtime records its own interrupt before issuing it. There are exactly two call sites in our runtime, the backstop and the preflight check, and both set a flag just before calling interrupt(). Simplified:
if (compactionRequested && pct >= cfg.backstopPct) {
compactionInterruptIssued = true; // recorded before the call
await activeQuery.interrupt();
}The flag travels with the query result into the reseed.
Second, when the flag is set, the ledger input relabels only the interrupt records at the very end of the transcript. An earlier, unrelated interrupt (a real person pressing stop an hour ago) is left alone. Simplified from our runtime:
const SDK_TOOL_REJECTION = "doesn't want to proceed with this tool use";
const SDK_INTERRUPT_MARKER = "[Request interrupted";
// Walk back from the end; stop at the first real user or assistant record.
function trailingInterruptIndices(records: TranscriptRecord[]): Set<number> {
const found = new Set<number>();
for (let i = records.length - 1; i >= 0; i--) {
const r = records[i];
if (r.type !== "user" && r.type !== "assistant") continue;
if (!isInterruptRecord(r)) break; // made only of SDK_TOOL_REJECTION / SDK_INTERRUPT_MARKER blocks
found.add(i);
}
return found;
}Those records become a system line: an automatic context refresh cut this call off, not the user; the call didn't complete; check its live effect, then re-run the step. "Check its live effect" matters. A download might have half-happened, and a file edit might have been written. The agent should look before it repeats anything.
Third, the extraction prompt got a rule: a system interrupt is never "waiting on the user", never a rejection and never a confirmation.
Seven ledger tests cover the relabel and the prompt rule, including one where an earlier interrupt has to stay exactly as it was. Three structural tests parse the runtime's source and fail if a third interrupt() call appears, if either call site stops setting the flag first, or if the flag stops reaching the reseed.
The relabel matches on the SDK's prose, because neither structured field is public and in our 0.3.282 transcripts the [Request interrupted marker carried no structured field at all. If Anthropic rewords that message, the match silently stops working and the agent goes back to blaming the user. Our tests use fixtures with today's wording copied in, so an SDK upgrade could break this without a single red test. That gap is the next thing to close. When a documented field marks a tool result as interrupted, I'll switch to it and delete the string matching.
The fix went out to all five of our hosts on October 8. The first backstop interrupt after that came on October 10. It cut 13 parallel tool calls at once. The ledger written under five minutes later recorded the cut as a system refresh, and it didn't list the step as refused by the user or as waiting on them.
If you run your own compaction
- Interrupt only as a backstop. Flag at a threshold and let the turn end on its own. Every interrupt you avoid is one fewer record you have to explain.
- Record your own interrupts before you issue them, and never infer who stopped a turn from transcript prose.
- Treat an interrupted tool as unknown, not as refused. Check the live state before re-running it.
- Don't read
getContextUsage()as the API's usage. Compare it with the last response's usage before you set thresholds tight. - If native summarization is good enough for your agent, use it. The
PreCompacthook (listed in our Claude Code hooks guide) lets you snapshot the transcript before it's summarized, which is far less work than owning the whole handoff.
The SDK side of this, the loop and its options, is in the Claude Agent SDK guide. If you work in Claude Code rather than the SDK, what survives compaction covers which instructions come back after /compact.
Common questions
What is context compaction?
Context compaction is what an agent does when its conversation approaches the model's context window: it replaces older history with something shorter so the session can continue. The Claude Agent SDK does this automatically by summarizing older history. Some runtimes, ours included, replace that with their own structured handoff.
Does interrupt() in the Claude Agent SDK record the user as the cause?
In our production SDK version (0.3.282), yes. The interrupted tool's result says "The user doesn't want to proceed with this tool use", and a marker begins "[Request interrupted by user". No documented field records whether a person or your own code issued the interrupt.
Should I build my own compaction?
Only if you've measured what the native summary loses for your agent. I did it because operations agents were losing exact IDs and rules across long sessions. If you own compaction, every interrupt you add is yours to explain, including this one.
Sources and measurement boundary
Event counts, timings and usage percentages come from Scalably's own runtime logs and transcripts for October 1 to 8, 2026, traced on one client's agents and reported here without client details. SDK behavior was read from the official Agent SDK documentation (TypeScript reference, Python reference, hooks and agent loop pages, checked October 10, 2026), the public TypeScript types of 0.3.282 and 0.3.296, and the SDK changelog through 0.3.296.