A large share of your context window can be gone before you type a word. Claude Code context window management is a budgeting problem, and most sessions run with no budget at all.
Frontend teams ship with a performance budget: every kilobyte of JavaScript has to justify itself. We apply the same discipline to tokens. This guide gives you the itemized bill, the profiler, and four levers that cut it.
What you’ll learn
- Read a
/contextbreakdown and separate standing overhead from conversation growth - Cut MCP overhead with tool search deferral, server pruning, and CLIs where they fit
- Steer Claude toward grep-first, range-limited file reads using prompts and hooks
- Delegate exploration to subagents that return fixed-price summaries
- Run a repeatable context budgeting workflow at the start of every serious session
Prerequisites
- A current Claude Code version: tool search deferral and the categorized
/contextoutput require recent builds - At least one MCP server configured, plus working knowledge of CLAUDE.md, skills, and hooks
- Comfort editing
.claude/settings.jsonand.mcp.jsonby hand
You never had the whole window
The window size is the gross figure. Before your first prompt, Claude Code loads the system prompt, system tool definitions, MCP tool names, custom agent descriptions, your CLAUDE.md files, auto memory, and the skill index. Auto-compaction also runs before the window is completely full, so the room you can actually work in is smaller than the headline number.
The official context window docs give representative startup costs:
- System prompt: ~4,200 tokens
- Project CLAUDE.md: ~1,800
- Global
~/.claude/CLAUDE.md: ~320 - Skill descriptions: ~450
- Auto memory: ~680
- Environment info: ~280
- Deferred MCP tool names: ~120
Anthropic labels these figures illustrative: your CLAUDE.md size and MCP setup change everything. The docs’ own framing is blunt: “Your prompt is tiny compared to what’s already loaded.”
Context window management is not only about running out of room. Anthropic’s engineering guidance calls context “a finite resource with diminishing marginal returns”: models have an attention budget, and recall degrades as the window fills. A fuller window is also a dumber window, so every token you cut buys quality, not just headroom.
Context window management starts with /context
Profile before you optimize. /context visualizes current usage as a colored grid with a breakdown by category, and adds optimization suggestions for context-heavy tools, memory bloat, and capacity warnings. Pass all to expand the per-item breakdown. The exact rendering changes between versions, so we don’t reproduce one here; run it in your own session.
Read the output as two kinds of line items. System tools, MCP tools, custom agents, and memory files are standing overhead: they ride along on every single API call. Messages are conversation growth, the only part that reflects work done. Whichever standing line is largest and within your control is where you cut first, and on setups with many MCP servers that is usually MCP.
Every token figure in this guide varies by version, model, and setup. The numbers from the official docs come from a simulation Anthropic labels illustrative. Profile your own session and verify against your environment.
Two companion commands round out the profiler. /memory opens the CLAUDE.md and auto memory files that loaded. On Pro, Max, Team, and Enterprise plans, /usage attributes recent consumption to skills, subagents, plugins, and individual MCP servers; press d or w to switch between the last 24 hours and the last 7 days (see the costs docs).
For longitudinal tracking beyond one session, instrument your Claude Code usage.
Add context usage to your status line so the number is always in view. Auto-compact mid-task is how budgets die: a visible gauge plus a /context check before each phase change means you choose when to compact, not the tool.
Lever 1: cut the MCP tax
With upfront loading, every tool definition on every connected server enters context at session start: names, descriptions, JSON schemas, and parameter descriptions, whether you call the tool or not. The totals grow fast. Before tool search existed, Scott Spence measured roughly 82k tokens of MCP tool definitions across his servers before any conversation.
Tool search defers most of it
Tool search is on by default in current builds. Only tool names and server instructions load at session start, and Claude fetches full schemas on demand. Control it with ENABLE_TOOL_SEARCH (see the MCP docs):
# Threshold mode: load deferrable tools upfront only while
# their definitions total less than 5% of the window
ENABLE_TOOL_SEARCH=auto:5 claude
# Legacy behavior: every schema upfront, no deferral
ENABLE_TOOL_SEARCH=false claude
Unset, all MCP tools are deferred. Claude Code falls back to upfront loading when ANTHROPIC_BASE_URL points at a non-first-party host (most proxies don’t forward tool_reference blocks), on Google Cloud Agent Platform models older than the Claude 4.5 generation, and on Microsoft Foundry deployments hosted on Azure. Tool search needs Sonnet 4.5, Haiku 4.5, Opus 4.5, or a later model. If your MCP line looks suspiciously large, check which endpoint you are on before blaming the servers.
Deferral reduces and relocates the cost; it does not erase it. Names and server instructions still load, each search is a tool-use round trip mid-task, and every tool that loads stays in context for the rest of the session.
To exempt a hot-path server from deferral, set alwaysLoad (added in v2.1.121):
{
"mcpServers": {
"sqlite": {
"command": "npx",
"args": ["-y", "mcp-sqlite", "./app.db"],
"alwaysLoad": true
}
}
}
Every tool from that server then loads at session start regardless of ENABLE_TOOL_SEARCH, and startup waits for the server’s tools, capped at the standard 5-second connect timeout. Keep it for the few tools Claude needs on every turn.
Prune and scope what you carry
Disable servers you are not using this session via /mcp. Scope is blast radius: a user-scoped server taxes every session in every repo, while a local- or project-scoped one loads only where it earns its place. For headless runs, --mcp-config plus --strict-mcp-config loads only the servers in the file you pass and ignores every other MCP configuration.
Prefer a CLI when one exists
The costs docs are direct: tools like gh, aws, gcloud, and sentry-cli are more context-efficient than MCP servers because they add no per-tool listing. Claude calls them through Bash and reads --help output only when it needs it.
Protocol is not the whole story, though. Mario Zechner benchmarked the same tool exposed as an MCP server and as a CLI and found overall cost roughly a wash; how well the tool was designed mattered more than the protocol. Keep the MCP server when it earns its rent: the service needs an OAuth flow a CLI cannot do, no CLI exists, or the environment has no shell. Otherwise make “use gh, not the GitHub MCP server” a line in your CLAUDE.md contract.
If you ship a server, you set the price
With schemas deferred, server instructions are the front door: Claude decides whether to search your server from tool names and instructions alone. The MCP docs ask server authors to state what category of tasks the tools handle, when Claude should search for them, and the key capabilities. Claude Code truncates each tool description and each server’s instructions at 2,048 characters by default, so put critical details first.
Consolidate related operations behind parameters instead of multiplying tools. Spence merged four web-search tools into one with a provider parameter and folded five Firecrawl modes into one tool; together with trimmed descriptions, that server went from 14,214 to 5,663 tokens with no features lost.
Response size is the other half of the tax. Claude Code warns when an MCP tool result exceeds 10,000 tokens and caps output at 25,000 by default. MAX_MCP_OUTPUT_TOKENS raises the cap, and a tool can declare its own limit with anthropic/maxResultSizeChars, but needing either usually means the tool should filter or paginate server-side. Finally, mark only your most-used tools as always loaded, with "anthropic/alwaysLoad": true in each tool’s _meta, and let tool search defer the rest.
Lever 2: read less
Once a session starts, file reads dominate context growth. In the official docs’ simulated session, four file reads plus one grep added ~7,500 tokens, roughly the entire startup overhead again. The fix lives in your prompt: name the file, the symbol, and the approximate location.
The token refresh bug is in the rotateToken function in
src/api/auth.ts, around line 140. Read just that function
and its call sites. Don't read the whole file.
The Read tool accepts offset and limit parameters, and a specific prompt like this steers Claude into searching for the symbol and reading only the relevant ranges. In the docs simulation, a whole-file read of auth.ts costs ~2,400 tokens; a range read of one function costs a fraction of that. Grep-before-read is the same discipline: a targeted search costs hundreds of tokens, while the candidate files it saves you from reading cost thousands.
The same logic scales up: pointing Claude at one package instead of the whole repo is the monorepo version of a range read. See scoping Claude Code in monorepos for that pattern.
Hooks are the force multiplier: they preprocess output before Claude ever sees it. Instead of Claude reading a 10,000-line test log to find failures, rewrite the command so only failures reach the context. This is the pattern from the costs docs:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": ".claude/hooks/filter-test-output.sh"
}
]
}
]
}
}
#!/bin/bash
# Rewrite test commands so Claude sees failures, not the
# thousands of lines of passing output.
input=$(cat)
cmd=$(echo "$input" | jq -r '.tool_input.command')
if [[ "$cmd" =~ ^(npm test|pytest|go test) ]]; then
filtered="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"
echo "$input" | jq --arg filtered "$filtered" '{
hookSpecificOutput: {
hookEventName: "PreToolUse",
permissionDecision: "allow",
updatedInput: (.tool_input + {command: $filtered})
}
}'
else
echo "{}"
fi
Now npm test returns each failure with five lines of context, capped at 100 lines total. The costs docs describe this kind of preprocessing as reducing context “from tens of thousands of tokens to hundreds.”
Lever 3: subagents as fixed-price exploration
Subagents convert open-ended exploration into a summary with a known price. In the official docs’ simulation, a research subagent read 6,100 tokens of files inside its own context window and returned a 420-token summary to the main session. In your terminal you see a brief notice that a subagent is working, then its result; its individual file reads never touch your context.
Use a subagent to research session timeout handling, then fix it.
Return file paths, line numbers, and a one-paragraph finding.
The savings live entirely in the return format. A subagent that answers with walls of text saves you nothing, so demand structure: file paths, line numbers, and a one-paragraph finding. The full firewall pattern, with return-format contracts and custom agent configs, is in subagents as context firewalls.
Subagents cut main-context spend but raise total token spend. Each one pays its own startup bill; in the docs simulation that is its own system prompt (~900 tokens), its own copy of project CLAUDE.md (~1,800), plus MCP and skill overhead (~970). Agent teams go further: the costs docs put them at roughly 7x the tokens of a standard session when teammates run in plan mode. Treat delegation as a context tactic first and a cost tactic only sometimes.
Two footprint reducers: the built-in Explore and Plan agents skip CLAUDE.md, and you can set model: haiku on custom agents for simple research tasks.
Lever 4: shrink the standing config
Every line of CLAUDE.md is charged on every session and re-injected from disk after every compaction. Official guidance: keep it under 200 lines and move workflow-specific instructions into skills, which load only when invoked. That split is the whole argument of CLAUDE.md as a contract, not a wiki.
Skills with disable-model-invocation: true are the extreme case: they cost zero context until you type the slash command. They do not even appear in the startup skill index. Here is the move for a PR-review workflow that used to live in CLAUDE.md:
# CLAUDE.md — always loaded, every session\n## PR review workflow (42 lines)\n1. Fetch the PR: gh pr view --json title,body,files\n2. Read the diff: gh pr diff\n3. Check CI: gh pr checks\n...39 more lines charged on every message, forever
# .claude/skills/review-pr/SKILL.md\n---\nname: review-pr\ndescription: Review a PR end to end\ndisable-model-invocation: true\n---\n# Loads only when you type /review-pr.\n# Zero tokens in every other session.
Path-scoped rules give you the same on-demand behavior for conventions. A rule with paths: frontmatter in .claude/rules/ loads only when Claude reads a matching file; in the docs simulation, a 380-token api-conventions.md loaded only when Claude touched src/api/.
Compaction changes the math of where instructions belong. After /compact, the system prompt, project-root CLAUDE.md, unscoped rules, auto memory, and the plan from plan mode reload from disk, and Claude Code re-reads up to five of the most recently modified files. Path-scoped rules and nested CLAUDE.md files are lost until re-triggered, and invoked skill bodies are re-injected capped at 5,000 tokens each.
Put must-survive instructions in the layers that reload. The full survival table, and how to steer what a compaction summary keeps, is in surviving auto-compact in Claude Code. If you are unsure whether an instruction belongs in CLAUDE.md, a skill, a rule, or a hook, use the decision matrix in choosing the right extension point.
A Claude Code context window management workflow
Prompt specificity is the zeroth tactic: “add input validation to the login function in auth.ts” beats “improve this codebase” before any tooling gets involved. On top of that, run this loop for every session that matters:
- 01
Baseline with /context at session start
Run
/contextbefore your first prompt and note the standing overhead. If more than a third of the window is gone before any work, fix levers 1 and 4 before starting: that overhead is billed on every message you send. - 02
Set a spend ceiling for exploration
Decide how many tokens the research phase gets, say 20k, and delegate anything bigger to a subagent. Entering plan mode first is the cheapest insurance against burning budget in the wrong direction; see plan mode: get plans worth approving.
- 03
Checkpoint before each big phase
Run
/contextagain before building, reviewing, or starting any long phase. If the messages category dominates, decide now, compact or clear, instead of letting auto-compact decide mid-task. - 04
Compact with focus, or clear at boundaries
Before a long new task in the same area, run
/compact focus on the auth bug fixso the summary keeps what you choose. Switching to unrelated work?/renamethe session,/clear, and/resumelater: stale context wastes tokens on every subsequent message. - 05
Re-profile and attribute
Compare the closing
/contextagainst your baseline. If the same category blows the budget every session, run/usageto attribute it to a specific server, skill, or subagent, then cut it at the source.
For genuinely large tasks, several current models support a 1 million token context window (via [1m] variants where the model offers one). Treat it as a bigger budget under the same discipline, not an escape hatch: the attention-budget problem scales with the window.
When we help teams roll Claude Code out, the standing overhead is where budgets quietly break: every developer inherits the same user-scoped MCP servers, the same bloated CLAUDE.md, and pays for them on every message. A shared baseline, meaning a vetted .mcp.json, a CLAUDE.md under 200 lines, and a /context check in onboarding, fixes that once for everyone instead of once per developer. That rollout groundwork is what our Enable engagements start with.
Next steps
- The escalation ladder for agentic tasks: match your context spend to task size instead of paying full price for small fixes
- Spec-first prompting for Claude Code: sharper prompts are the zeroth budgeting tactic; here is how to write them
- Subagents as context firewalls: the return-format contracts behind Lever 3
External references worth bookmarking: the official context window documentation with its interactive cost simulation, the Claude Code MCP docs for the tool search reference, and Anthropic’s effective context engineering post on attention budgets and context rot.