Claude Code subagent context isolation

Use Claude Code subagents as context firewalls. Three patterns that keep test output, search noise, and review bias out of your main thread.

on this page

Your main thread’s context window is the scarcest resource in a Claude Code session. Claude Code subagent context isolation is how you defend it. A subagent can burn 100,000 tokens grepping, reading files, and running tests. The only thing that crosses back into your conversation is its final message.

Most writeups sell subagents as a parallelism feature. That framing is popular, and we think it is the wrong lead. Parallelism is a speed bonus. Isolation is the structural win, and it pays off even when you run one subagent at a time.

What you’ll learn

  • Trace exactly what crosses the subagent boundary, in both directions
  • Apply three firewall patterns: fan-out search, adversarial verification, and noisy-op quarantine
  • Design output contracts that stop subagent results from re-polluting your main thread
  • Decide when a forked subagent, /btw, or the main thread beats a fresh subagent
  • Build a read-only adversarial verifier agent you invoke with one @-mention

Prerequisites

  • A current Claude Code v2.1.x release (v2.1.219 or later for the nesting default described below; the Task tool became the Agent tool in v2.1.63, and Task(...) permission rules still work as aliases)
  • Daily, hands-on Claude Code use — this guide skips setup and definitions
  • A project with a real test suite, so you can reproduce the quarantine pattern
  • Working knowledge of .claude/ configuration files

The wrong reason to use subagents

Anthropic’s engineering post on its multi-agent research system contains the number that should reorganize how you think about subagents. Token usage alone explained 80% of performance variance on their research evals. Model choice and tool-call count explained most of the rest.

Their multi-agent setup — an Opus lead with Sonnet subagents — beat single-agent Opus by 90.2% on that internal research eval. That figure comes from a research benchmark, not coding tasks, so do not assume it transfers directly. The mechanism transfers, though. In Anthropic’s words, subagents operate “in parallel with their own context windows… before condensing the most important tokens for the lead research agent.”

Read that as compression, not concurrency. Each subagent spends tokens in a window that never counts against your orchestrating thread. A subagent is a firewall with a narrow gate: arbitrary work happens on the far side, and only a condensed conclusion passes through.

That makes subagents a spending instrument inside your larger context budget in Claude Code. You are not saving tokens overall — you are choosing which window absorbs them. The rest of this guide treats every pattern in those terms: tokens kept out of the main thread, not wall-clock time saved.

How subagent context isolation works

The boundary is precisely documented, and it is asymmetric. You need both directions clear before the patterns make sense.

Inbound. A subagent starts in a fresh, isolated context window. It does not see your conversation history, the files you have read, your invoked skills, or the session’s auto memory. It receives five things at startup:

  • Its own system prompt (the agent file’s markdown body)
  • The delegation task message Claude writes
  • CLAUDE.md plus the full memory hierarchy
  • A git status snapshot
  • Any preloaded skills
Note

The built-in Explore and Plan agents skip CLAUDE.md and the git snapshot entirely to stay cheap. A custom agent can skip CLAUDE.md too, with omitClaudeMd: true in its frontmatter. If a project rule matters to the delegated work — “ignore vendor/”, “never touch generated files” — restate it in the delegation prompt.

Outbound. Only the subagent’s final message enters your context, plus a small metadata trailer with token counts and duration. Every intermediate tool call, file read, and line of test output stays in the subagent’s own transcript. Claude Code stores it separately at ~/.claude/projects/{project}/{sessionId}/subagents/agent-{agentId}.jsonl.

The official context-window walkthrough puts illustrative numbers on this. In its simulated session, a subagent reads 6,100 tokens of files and the main thread receives a 420-token summary. The docs’ own caption: “That’s the context savings.”

Two properties round out the mechanics. Subagents auto-compact with the same logic as the main thread, so a 100k-token exploration does not die at the window edge.

Subagent transcripts are also untouched when your main conversation compacts. The built-in Plan agent uses exactly this to keep plan mode’s research out of your window while the plan itself survives.

When you need to map several independent subsystems before a cross-cutting change, fan the exploration out. The official docs’ example prompt:

Research the authentication, database, and API modules in parallel
using separate subagents

Each subagent explores its slice in its own window, and Claude synthesizes the condensed reports. The parallelism is nice. The isolation is the point: three exploration transcripts’ worth of file contents never enter your thread.

The docs attach a warning worth quoting verbatim: “Running many subagents that each return detailed results can consume significant context.” The firewall only works if the gate stays narrow. That means you firewall the output contract, not just the work — specify the shape of the answer, not merely the task.

Before
Research the authentication, database, and API modules in parallel using separate subagents
After
Fan out three Explore subagents: auth, database, and API modules. Each returns at most 10 bullet points: entry points, key abstractions, and anything that touches session state. Nothing else.

The first prompt is the docs’ verbatim pattern and it works, but each agent decides how much to return. The second caps the total return payload at roughly thirty bullets, and names exactly what you care about. We default to the constrained phrasing, which is consistent with the docs’ guidance — test it against your own repo, since return discipline varies with the task.

Fan-out fits independent research paths. If step two depends on step one’s full findings, chain the work instead — or keep it in the main thread, as covered below.

Pattern 2: adversarial verification

The inbound side of the firewall blocks something more valuable than tokens: bias. When the same model reviews its own work in the same context window, shared context preserves every generation-time assumption. You get agreement, not review.

A fresh-context subagent reviewer is structurally unable to rubber-stamp reasoning it never saw. The claude.com blog on subagents makes the case directly: a subagent gets “a clean slate because it doesn’t inherit the assumptions, context, or blind spots from the primary conversation.”

Design the delegation accordingly. The verifier receives the claim (“the token rotation fix resolves the 401-after-refresh bug; auth tests pass”) and the artifacts (the diff, file paths). It never receives the parent’s reasoning.

Restrict it to read-only tools: Read, Grep, Glob, Bash with no Edit or Write, matching the docs’ own code-reviewer example. The verifier can inspect and re-run tests, but it cannot “fix” its way to a pass.

This is one loop in a broader discipline of verification loops for agentic code, and it is the machine half of a pair. The human half is reviewing AI diffs without rubber-stamping — the subagent verifier catches what shared context hides, and you catch what both models miss. The full agent file is in the build section below.

Pattern 3: quarantine noisy operations

The docs’ first “common pattern” is the everyday one: isolate high-volume operations. Test runs, documentation fetches, and log digging consume context you will never reference again. Delegate them, and the verbose output stays in the subagent’s transcript.

The official prompt, verbatim:

Use a subagent to run the test suite and report only the failing
tests with their error messages

Notice the prompt carries its own output contract: “report only the failing tests with their error messages.” What comes back is a short report: a pass/fail count, each failing test with its error, and perhaps a line on what the failures have in common. The thousands of tokens of raw test-runner output exist only in the subagent transcript. Your thread pays for the report, not the run.

Pro tip

Set model: haiku on agents that do noisy grunt work. Running tests and summarizing failures is grep-and-condense work that does not need the frontier model, and it directly cuts the token-cost multiplier that makes firewalling expensive.

Forks: the one-way firewall

Named subagents give you a two-way firewall: fresh context in, summary out. A forked subagent deliberately gives you half of that. A fork inherits your entire conversation — full history, every file read — so input isolation is gone. Output isolation remains: the fork’s tool calls stay out of your thread, and only its final result comes back as a message in your main conversation.

You start one yourself with /subtask followed by the task (v2.1.212 or later), and in interactive sessions Claude can spawn forks on its own. Do not confuse this with the current /fork command. /fork now copies the whole conversation into a separate background session that you keep working alongside, and nothing flows back into your thread. On v2.1.161 through v2.1.211, and whenever agent view is turned off, /fork still starts a forked subagent instead.

Forks are also cheaper than fresh subagents because they reuse the parent’s prompt cache. Reach for one when the side task needs everything you have already established. “Given all of this, check whether the migration plan breaks the staging config” is fork-shaped work.

Reach for a named subagent when inherited context would hurt — verification being the sharpest case. A forked verifier inherits the implementation journey and its blind spots with it. Our rule of thumb: fork for informed side quests, fresh subagent for independent judgment.

When the firewall costs you

Subagent context isolation is a deliberate spend, not a free lunch. Anthropic’s own numbers: multi-agent systems used about 15x more tokens than chat, and single agents about 4x. Their guidance is blunt — multi-agent work pays off for high-value tasks that exceed one context window or parallelize cleanly, and it is a poor fit otherwise.

Keep work in the main thread when:

  • You are iterating. Frequent back-and-forth or refinement loops make cold-start subagents slow and wasteful.
  • Phases share context. Plan, build, and test on the same code want one window, not three firewalled ones.
  • Step two needs step one’s full output. The firewall strips exactly what you need. A condensed summary of a refactor is not enough to continue the refactor.
  • The change is small. Delegation overhead swamps a two-file edit.
  • Latency matters. Subagents start fresh and spend time gathering context you already have.

Two more anti-patterns: never point multiple subagents at the same file, and do not use subagents for sustained cross-agent coordination — agent teams are the tool for that. Where a task sits on this spectrum is an escalation-ladder decision: main thread, fork, subagent, then teams.

For a quick side question about something already in context, /btw beats all of them — full context, no tools, and the answer is discarded. And remember the second-order benefit when you do firewall: every token you keep out of the main thread delays compaction, which is half the battle of surviving auto-compact.

Build your own firewall agent

Custom agents live as markdown files with YAML frontmatter in .claude/agents/ (project) or ~/.claude/agents/ (user). Only name and description are required, but the firewall-relevant fields are tools, disallowedTools, model, maxTurns, background, isolation, and mcpServers.

Here is the adversarial verifier from Pattern 2 as a complete agent file:

---
name: adversarial-verifier
description: Independently verifies claimed fixes. Use after
  implementation, before commit. Gets the claim and the diff,
  never the reasoning.
tools: Read, Grep, Glob, Bash
model: inherit
---

You are an adversarial verifier. You receive a claim about what a
change does and the changed files. You did not write this code and
you do not trust the claim.

1. Run git diff to see the actual change
2. Try to falsify the claim: edge cases, error paths, concurrent
   access
3. Run the relevant tests yourself; do not trust reported results
4. Return a verdict: CONFIRMED or REFUTED, with evidence, in under
   300 words

The frontmatter format and read-only tool restriction follow the docs’ code-reviewer example. The 300-word verdict cap is the output contract, written into the agent itself so every invocation inherits it.

  1. 01

    Create the agent file

    Save the file above as .claude/agents/adversarial-verifier.md. Project-level agents ship with the repo, so your whole team gets the same verifier.

  2. 02

    Let Claude Code pick it up

    Claude Code watches .claude/agents/ and ~/.claude/agents/ and picks up new or edited agent files within a few seconds, with no restart. Three cases still need one: the agents directory did not exist when the session started, the agent lives in a directory added with --add-dir or /add-dir, or the session was started with --disable-slash-commands.

  3. 03

    Invoke it with an @-mention

    An @-mention guarantees this subagent runs; Claude still writes the task prompt. For example:

    @agent-adversarial-verifier the token rotation fix in
    src/lib/tokens.ts resolves the 401-after-refresh bug and all
    auth tests pass
  4. 04

    Read the verdict, not the transcript

    Only the CONFIRMED/REFUTED verdict and its evidence enter your thread. If you need the full investigation, the subagent transcript is on disk under ~/.claude/projects/.

Three frontmatter fields extend the firewall beyond context. Setting mcpServers scopes an MCP server to the subagent, so its tool schemas never enter your parent window at all — a direct answer to the MCP line item in your context budget. Setting isolation: worktree adds filesystem isolation on top of context isolation, the same mechanism behind parallel worktree workflows. And background: true keeps the agent in the background; background subagents surface their permission prompts in your main session.

Subagents can also spawn their own subagents. Since v2.1.219 the default limit is three layers below your main conversation, and CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH changes it. (v2.1.172 through v2.1.216 allowed a fixed five layers.) In an interactive session, only the top-level subagent’s summary returns to your main thread. The firewall composes.

On a team, the firewall only holds if everyone uses the same one. A verifier in one engineer’s ~/.claude/agents/ protects one engineer; the same file committed to .claude/agents/, with a written output contract, becomes a shared standard for what “verified” means. Getting those shared agents and contracts in place is part of how we help with rolling Claude Code out across engineering teams.


Next steps

Related service

This is part of our Enable work — make the team able to operate it.