Claude Code auto-compact: survive the cliff

Learn what Claude Code auto-compact actually drops, why sessions degrade after compaction, and the disk-first workflow that makes resets free.

on this page

What you’ll learn

  • Identify exactly what the compaction summary keeps and what it silently drops
  • Read the official survival table and promote must-survive rules to the right file
  • Pick your own compaction boundary with /context instead of trusting a threshold number
  • Steer summaries with /compact focus instructions and a Compact Instructions section
  • Externalize session state into plan files and handoff notes so any context window is disposable

Prerequisites

  • Working fluency with Claude Code: CLAUDE.md, rules, skills, plan mode, and subagents
  • A project where you run long, multi-step agentic sessions
  • A recent Claude Code CLI version (compaction behavior shifts between releases)

Picture a session that opens with a single standing rule, stated in chat: never spawn parallel subagents. Hours in, auto-compact fires mid-task, and the summary paraphrases the constraint away or drops it. The next action fans out a batch of subagents at once, and the usage budget goes with it.

That is the compaction cliff. Claude Code auto-compact fires when the context window fills, not when your task sits at a clean boundary. This guide covers what the summary keeps, what it silently drops, and how to work so that losing any single context window costs you nothing.

How Claude Code auto-compact degrades sessions

Quality does not drop in one step. It drops in two.

First, context rot. Anthropic’s engineering guidance is blunt: as the token count in the window grows, the model’s ability to accurately recall information from it decreases. Your session is already sliding before compaction fires.

Second, the lossy summary. When Claude Code auto-compact triggers, it condenses the conversation into a structured summary written by the model itself. Whatever imprecision context rot introduced gets locked into that paraphrase.

You can switch the automatic pass off — autoCompactEnabled: false in settings, or DISABLE_AUTO_COMPACT=1 for one session — but that does not remove the cliff. The window still fills, and you still have to compact or clear before you hit the limit. We would rather stop fighting the mechanism and make it irrelevant instead.

What the summary keeps and drops

The docs are specific about the post-compact state. The summary preserves your requests and intent, key technical concepts, files examined or modified with important snippets, errors and how they were fixed, pending tasks, and current work.

It replaces the verbatim conversation. Full tool outputs and intermediate reasoning are gone. Claude Code then re-reads up to five of the files Claude read or edited, most recently modified first; a file over 5,000 tokens comes back only as a path reference. Everything else Claude read is available only as a mention in the summary.

The practical translation: most of what Claude read is gone, apart from a handful of recent files, and everything Claude concluded survives only as paraphrase. The precision casualties are predictable:

  • Exact error messages and stack traces
  • Exact identifiers: variable names, table names, flag values
  • Constraints you stated mid-conversation (“never spawn parallel subagents”)
  • “We tried X, it failed because Y” reasoning that prevents repeated mistakes
  • Detailed instructions from early in the conversation

Claude Code auto-compact also runs in two passes. It clears older tool outputs first, then summarizes the conversation if space is still tight. Either pass can take something you were relying on.

What survives compaction, mechanism by mechanism

The official context-window docs publish a survival table. It is the mechanical answer to “why did Claude’s behavior change after compaction”, and most practitioner posts do not know it exists.

MechanismAfter compaction
System prompt and output styleBoth still apply
Project-root CLAUDE.md and unscoped rulesRe-injected from disk
Auto memory (MEMORY.md)Re-injected from disk
Git status snapshotA fresh one is read from the repository
The plan Claude wrote in plan modeRe-injected from disk
Rules with paths: frontmatterLost until a matching file is read again
Nested CLAUDE.md in subdirectoriesLost until a file in that directory is read again
Files Claude read or editedUp to five re-read, most recently modified first
Invoked skill bodiesRe-injected; capped at 5,000 tokens per skill, 25,000 total, oldest dropped first
Skill listing (descriptions)Not re-injected after /compact
Context that hooks added earlierSummarized with the rest of the conversation
SessionStart hooks matching the compact sourceRun again; their output is added to the compacted context
Warning

Path-scoped rules and nested CLAUDE.md files load into message history when their trigger file is read, so compaction summarizes them away with everything else. They stay dormant until a matching file is read again. This is the most common cause of the “Claude forgot my rule” mystery after a long session.

The fix is placement, not repetition. If a rule must survive compaction, drop its paths: frontmatter or move it into the project-root CLAUDE.md, which re-enters from disk after every compaction. This is the operational argument for treating CLAUDE.md as a contract, not a wiki: the contract persists because it lives on disk, not in chat history.

Skills carry two traps of their own. The skill listing is not re-injected after /compact, so only skills you actually invoked stay available in context. And re-injected skill bodies are truncated to their first 5,000 tokens — put your most important instructions at the top of SKILL.md.

Hooks sit outside the context window. The hook itself keeps running after compaction because it is code, not context, which makes hooks the right home for rules that must never lapse. What a hook added to the conversation earlier gets summarized like everything else. A SessionStart hook that matches the compact source runs again after compaction, which makes it a way to re-inject context you choose.

Know where you are with /context

Do not memorize a trigger percentage. Where the automatic pass fires depends on your model, your provider, and your configuration, and it has shifted across releases. You can move it earlier yourself: /autocompact 500k (v2.1.221 or later) sets how full the window gets before auto-compact runs, and CLAUDE_AUTOCOMPACT_PCT_OVERRIDE can lower the trigger percentage but never raise it.

Note

The docs publish the current default thresholds per model on the model configuration page, and they change between releases. The docs also note a thrashing guard: if context refills immediately after each summary, Claude Code stops auto-compacting after a few attempts and shows an error instead of looping.

Observe instead. Run /context for a live picture of what is using the window, shown as a colored grid with optimization suggestions for context-heavy tools, memory bloat, and capacity warnings. Check it at natural pauses rather than waiting for the automatic pass to tell you.

Deliberate context budgeting keeps you off the cliff in the first place, and it starts with knowing what eats your window before you type anything, from CLAUDE.md and memory to MCP tool definitions.

The zone model we use for context usage:

  • Below 30% used: keep working.
  • Between 30% and 50%: start watching for clean boundaries.
  • Between 50% and 60%: finish the current micro-task, then compact or clear on your own terms.
  • Past 60%: reset before starting anything major.

Delegation stretches every zone. Route large file reads and log dumps through subagents acting as context firewalls, so raw contents never enter the main window — and never get summarized out of it.

Steer the summary when you do compact

When the next task continues the current work, a steered /compact beats both auto-compact and a full /clear. The command accepts instructions, and the summary keeps what you choose instead of what the automatic pass guesses is important.

Tell it exactly what to keep verbatim and what to discard:

/compact Focus on the auth bug fix. KEEP verbatim: the acceptance
criteria, the list of files changed (src/auth/session.ts,
src/auth/refresh.ts, middleware/verify.ts), the decision to rotate
refresh tokens server-side, and the two open test failures with
their exact error text. DISCARD: file contents we read, npm install
output, and the abandoned cookie-based approach.

The summarization itself does not appear in your terminal. Per the docs, you see a “Conversation compacted” message and a one-line note for each file Claude Code re-read, then continue with a window holding startup content, the steered summary, and those re-read files.

You cannot schedule the automatic pass, but you can shape it. Add a “Compact Instructions” section to your project-root CLAUDE.md and both manual and automatic compaction will follow it:

## Compact Instructions

When compacting this conversation, always preserve:
- The current goal and its acceptance criteria, verbatim
- Exact file paths changed this session, with one line on why
- Decisions made and alternatives rejected, with reasons
- Open bugs and failing tests, with exact error text
- The numbered list of remaining steps

Discard file contents that can be re-read from disk, dependency
install output, and approaches we explicitly abandoned.

This section is compaction insurance you write once. It costs a few hundred tokens of CLAUDE.md space and pays out every time the automatic pass fires at a boundary you did not choose.

Externalize state so compaction stops mattering

Steered summaries reduce the damage. Externalized state removes it. Anthropic’s engineering team calls the pattern structured note-taking: the agent writes notes to persistent storage outside the context window and pulls them back in later.

For Claude Code sessions, that converges on three artifacts:

  • A plan file with numbered steps and checkboxes, updated as steps complete. Plan mode, used well, produces this file for you at the start of the task, and the plan it writes is re-injected from disk after every compaction.
  • A handoff note written before a reset: goal, repo state, decisions with reasons, open bugs with exact errors, next steps.
  • A rehydration ritual: explicitly re-read both files after any compact or clear.

A handoff note looks like this:

# Handoff: billing migration

## Goal
Move charge dedup from app-level checks to Postgres advisory
locks. Done when TestChargeIdempotency passes 50 consecutive runs.

## State
Branch: feat/billing-advisory-locks (3 commits ahead of main)
Changed: internal/billing/migrate.go, internal/billing/locks.go,
internal/billing/billing_test.go

## Decisions
- Advisory locks over SELECT FOR UPDATE: avoids lock queuing on
  the hot charges table under retry storms.

## Open bugs
- TestChargeIdempotency flakes with: duplicate key value violates
  unique constraint "charges_pkey"

## Next steps
1. Add the idempotency-key check before insert
2. Re-run the flake loop: go test -run TestChargeIdempotency -count=50
3. Wire retry backoff into the API layer

The full procedure, from a busy session to a clean one:

  1. 01

    Write the handoff before you stop

    Have Claude write the note while full context still exists. Be explicit that the audience is a fresh session with zero context:

    We're at ~60% context and about to switch from the migration
    script to the API layer. Before we stop: write
    docs/handoff/2026-07-02-billing-migration.md containing:
    1. The goal and acceptance criteria for this migration
    2. Exact files changed so far and why (paths, not summaries)
    3. Decisions we made and rejected alternatives (include the
       reason we chose advisory locks over SELECT FOR UPDATE)
    4. Open bugs: the flaky test in billing_test.go and its exact
       error message
    5. Next steps as a numbered list, starting with the
       idempotency check
    Quote exact error messages and identifiers verbatim — this
    file is for a fresh session with zero context.
  2. 02

    Verify it quotes exact strings

    Open the file yourself. Check that error messages, file paths, and identifiers appear verbatim, not paraphrased. A handoff full of summaries recreates the exact loss you are trying to avoid.

  3. 03

    Clear the session

    Run /clear. You are not losing anything — the state that matters is now on disk, and the degraded conversation history was a liability, not an asset.

  4. 04

    Rehydrate from disk

    In the fresh session, load only the two files and pin the settled decisions:

    Read docs/handoff/2026-07-02-billing-migration.md and
    docs/plans/billing-migration.md. Continue from step 3 of the
    next steps. Do not re-derive decisions already recorded in
    the handoff — treat them as settled.
  5. 05

    Confirm before any edits

    Ask Claude to restate the next step and its acceptance criteria before it touches code. If the restatement is wrong, fix the handoff file, not the conversation.

One trap in the ritual: after an auto-compact, the summary may reference your handoff note and your own plan files, but only the plan-mode plan is guaranteed to come back. Other files return only if they are among the five most recently modified that Claude Code re-reads, and anything over 5,000 tokens comes back as a bare path. Re-read your handoff and plan files explicitly before continuing — a mention in a summary is not the file.

When to /clear instead of limping on

The compact-versus-clear call comes down to a simple decision rule.

Use a steered /compact when you hit a phase boundary and the next task continues the current work. The summary carries forward genuinely useful state, and a full reset would throw it away.

Use /clear plus a tight brief when you switch to unrelated work, or when the session already auto-compacted once and quality is visibly off. It is also the right call when correcting drift would cost more tokens than re-establishing context. That last case is the correction tax in its purest form — every fix to a confused session costs more than a clean restart would.

A tight restart brief is a spec, and it benefits from the same discipline as spec-first prompting. It contains the goal, the acceptance criteria, pointers to the handoff and plan files, and the single next step. Nothing else — the fresh window’s emptiness is the point.

Pro tip

Compact at boundaries you choose. If Claude Code auto-compact fires mid-task anyway, treat the session as suspect: before letting the next action run, verify it against your plan file. The runaway-subagents scenario at the top of this guide is what happens when nobody checks the summary against the standing rules.

On a team, the cost of the cliff multiplies: every engineer loses a rule at a different moment, and nobody notices the same way twice. The fix scales because it is placement, not vigilance — must-survive rules in the project-root CLAUDE.md or a hook, a shared Compact Instructions section, and a handoff format everyone uses. Settling those conventions once for every engineer is part of how we approach rolling Claude Code out across engineering teams.

Next steps

Related service

This is part of our Enable work — make the team able to operate it.