Plan mode vs subagents: pick the right rung

Decide when to use plan mode vs subagents in Claude Code with an escalation ladder that prices every rung, from hand edits to headless fan-out.

on this page

What you’ll learn

  • Match any task to one of seven escalation rungs, from hand edits to headless fan-out
  • Apply the documented triggers for when to use plan mode vs subagents in Claude Code
  • Price over-escalation per rung: plan-mode overhead, subagent cold starts, token multiples
  • Recognize when the right move is climbing down to a deterministic tool
  • Use your correction count as the feedback signal that recalibrates future picks

Prerequisites

  • Plan mode, subagents, and claude -p already part of your workflow — this guide covers when, not how
  • Claude Code on a current release, with permission rules you understand
  • Familiarity with git worktrees for the parallel-session rung
  • A real backlog to calibrate against, not toy tasks

The calibration problem

Every experienced Claude Code user has a favorite hammer.

Some open plan mode for everything. Some spawn subagents for a two-file read. Some one-shot a cross-cutting migration and spend the afternoon correcting it.

The skill that separates power users is not knowing that these modes exist — you already know that. The skill is calibration: matching the amount of agentic machinery to the task in front of you. Knowing when to use plan mode vs subagents in Claude Code is one comparison inside a larger decision.

Miscalibration costs you in both directions. Over-escalate and you pay setup overhead, cold-start context, and token multiples for a task a one-liner would handle. Under-escalate and you pay in corrections, polluted context, and rework — the more expensive failure.

One constraint drives the whole ladder. As the official best-practices docs put it: “Claude’s context window fills up fast, and performance degrades as it fills.” Every rung is a different answer to that constraint, which is why context budgeting sits underneath all of them.

The docs describe their own verification sequence as one where “each step trades setup for attention.” That is the currency on every rung below: you spend setup time to buy the right to look away.

The ladder at a glance

Each rung adds setup cost and autonomy. A specific, checkable task property justifies each one — not how impressive it feels.

RungEscalate here when…Cost of using it too early
0 — No agentA script, linter, or hook already does it deterministicallyTokens and minutes for what a tool does in seconds
1 — One-shot promptYou can describe the diff in one sentenceNone; this is the floor
2 — One-shot + checkSame scope, but you want to walk awayA slightly longer prompt
3 — Plan mode3+ files, architectural choice, or unfamiliar codePlanning overhead on tasks with obvious diffs
4 — SubagentsExploration hits 10+ files, or 3+ independent piecesCold context, several times the token cost
5 — Parallel sessionsIndependent workstreams, or writer/reviewer separationCoordination overhead, merge conflicts
6 — HeadlessRecurring or batch, with verifiable outputUnverifiable unattended runs

Screenshot this table if you want. The rest of the guide is what makes it usable: the trigger and the over-escalation cost for each rung, with the prompts that belong there.

Rung 0: route work away from the agent

The most underrated rung is the one below the ladder. If a check is mechanical, the change is fully known, or a reliable tool already exists, the agent is the wrong tool. Our rule: enforce anything mechanically checkable with hooks or scripts, not by instructing the LLM.

An agent can spend twenty minutes rediscovering what a static analyzer reports in seconds, non-deterministically and at token cost. If the existing tool works reliably, the burden of proof is on the agent replacing it.

This is why mechanical rules belong in hooks, as guardrails rather than suggestions. A formatter in a PostToolUse hook fires every time. The same rule written in CLAUDE.md is a suggestion the model can miss under context pressure.

Choosing rung 0 is a power-user move, not a failure to be agentic. The top of the calibration skill is knowing when to climb down.

Rungs 1 and 2: the one-shot tier

The floor of agentic work is a single, well-scoped prompt in your main session. The official docs give the cleanest trigger you will find. For tasks where the scope is clear and the fix is small — a typo, a log line, a rename — ask Claude to do it directly.

Tip

The single most quotable heuristic in the official docs: “If you could describe the diff in one sentence, skip the plan.” Say the sentence, watch the edit, move on.

What makes a one-shot succeed is precision, not luck. Name the files, state the constraint, define done. That discipline is its own skill — spec-first prompting covers how to write the sentence so it lands on the first pass.

Rung 2 is the cheapest escalation on the ladder: same task scope, plus a check Claude can run itself. Append “run the test suite and fix any failures” or “build and confirm zero type errors” to the prompt. That one clause converts a watched session into one you can leave, because the agent now has a pass/fail signal instead of your eyeballs.

If the task has no runnable check, build one before escalating further. Verification loops shows how to give every task a machine-checkable definition of done — you will need it again at rung 6.

Rung 3: plan first

Escalate to plan mode when any of these hold: the change touches multiple files (we draw the line at three), there is a real architectural choice to make, or the code is unfamiliar. The fourth trigger is the inverse of rung 1 — you cannot describe the diff in one sentence. The docs are explicit that planning is most useful when you are uncertain about the approach — and equally explicit that it “adds overhead”.

The canonical rung-3 sequence from the official docs looks like this:

  1. 01

    Explore in plan mode

    Cycle into plan mode with Shift+Tab and point Claude at the relevant code. No edits can happen yet, so exploration is free of risk:

    read /src/auth and understand how we handle sessions and login.
    also look at how we manage environment variables for secrets.
  2. 02

    Ask for the plan

    Still in plan mode, ask for the plan against what it just read:

    I want to add Google OAuth. What files need to change?
    What's the session flow? Create a plan.
  3. 03

    Review before approving

    Read the plan as a reviewer, not a spectator. Press Ctrl+G to open it in your editor and cut anything out of scope. An approved plan is a contract for the execution phase.

  4. 04

    Execute against the approved plan

    Exit plan mode and hand over execution with the verification clause from rung 2 built in:

    implement the OAuth flow from your plan. write tests for the
    callback handler, run the test suite and fix any failures.

Expected behavior: in plan mode Claude reads files and produces a plan for approval without editing anything; after approval it edits, runs the tests, and iterates on failures.

The step most people skip is step 3, and it is the entire point of the rung. If you approve plans without reading them, you are paying plan-mode overhead for rung-1 behavior. Plan mode: get plans worth approving covers how to make that review fast and real.

When to use plan mode vs subagents

These two rungs blur together because both feel like “the careful option.” They solve different problems.

Plan mode manages uncertainty inside one workstream. One context window, one line of work — you are buying a checkpoint where a human approves the approach before edits begin.

Subagents manage context volume and independence, and Anthropic’s post on how and when to use subagents gives the crisp trigger. A task that requires exploring ten or more files, or three or more independent pieces of work, is a strong signal to reach for subagents. The exploration would poison your main context; the independent pieces do not need each other’s intermediate output.

The same post lists the anti-patterns worth memorizing:

  • Sequential dependent work, where step two needs step one’s full output
  • Parallel edits to the same file
  • Small focused tasks, where the overhead is not justified
  • A sprawling roster of specialist agents, which makes automatic delegation unreliable

The two rungs also compose. A common advanced pattern is subagent exploration feeding a plan: delegate the wide read, then plan in your still-clean main session.

Rung 4: subagents

Subagents exist to solve a context problem, not to win a parallelism trophy. Each one gets a fresh context window, does its work there, and returns only a summary — which is why they work as context firewalls for your main session.

The first canonical use is wide investigation. This prompt is verbatim from the official docs:

Output

❯ Use subagents to investigate how our authentication system handles token refresh, and whether we have any existing OAuth utilities I should reuse.

Expected behavior: Claude spawns an exploration subagent, and your terminal shows a brief notice that it is working rather than its individual file reads. It returns a findings paragraph — not raw file contents — and your main context stays clean for the actual work.

The second canonical use is adversarial review in a fresh context. The reviewer sees only the diff and the criteria, never the reasoning that produced the change, so it grades the result on its own terms:

Output

❯ Use a subagent to review the rate limiter diff against PLAN.md. Check that every requirement is implemented, the listed edge cases have tests, and nothing outside the task’s scope changed. Report gaps, not style preferences.

Note the last clause. The docs warn that a reviewer prompted to find gaps will report some even when the work is sound. Scope it to correctness or you will over-engineer on its feedback.

Pair this rung with reviewing AI diffs without rubber-stamping — the subagent reviews, but you still decide.

Warning

Subagents start with zero task context and bill for every window they open: each loads its own system prompt, CLAUDE.md, and tool overhead before doing any work. The costs docs put agent teams at roughly 7x the tokens of a standard session when teammates run in plan mode. You are trading tokens for a clean main context and wall-clock time on parallelizable work.

For a quick targeted edit, the main conversation is strictly faster and cheaper. Do not spawn a subagent for a two-file read.

Rung 5: parallel sessions

When you have multiple genuinely independent workstreams — not independent subtasks of one feature, but separate features — one session stops being the bottleneck you can fix with subagents. Run separate Claude Code sessions in parallel git worktrees, one branch and one working directory each.

The trigger is independence plus duration. Two features that will each take an hour and never touch the same files justify the setup. Two tasks that overlap in src/api/ do not; you will pay the coordination cost back in merge conflicts.

The other rung-5 pattern is role separation: a writer session and a reviewer session on the same change, in different worktrees. You get rung-4’s fresh-context review with a human-speed interface. Agent teams extend this rung for longer autonomous runs, but treat them as an experiment, not a default.

Rung 6: headless runs

The top rung is claude -p: non-interactive execution for CI, cron, pre-commit, and batch fan-out. It is also the rung with hard prerequisites, because nobody is watching. Go headless only when the task is recurring or batch, the failure modes are understood, and the output is verifiable after the fact.

A prompt earns rung 6 by proving itself on the lower rungs first. The docs’ fan-out recipe builds that in — test on a few files, refine the prompt based on what goes wrong with the first two or three, then run at scale:

for file in $(cat files.txt); do
  claude -p "Migrate $file from Python 2 to Python 3. Return OK or FAIL." \
    --allowedTools "Edit,Bash(git commit *)"
done

Two details in that loop do the safety work. --allowedTools scopes what an unattended run may do, and “Return OK or FAIL” makes every result machine-checkable — rung 2’s verification clause, industrialized.

The permission syntax uses prefix matching, and the space before * matters: Bash(git diff *) matches git diff --stat but not git diff-index. Here is the docs’ single-shot example, a commit generated from staged changes:

claude -p "Look at my staged changes and create an appropriate commit" \
  --allowedTools "Bash(git diff *),Bash(git log *),Bash(git status *),Bash(git commit *)"

For CI, add --bare to skip auto-discovery of hooks, skills, subagents, plugins, MCP servers, auto memory, and CLAUDE.md. The headless docs call it the recommended mode for scripted calls, and it makes runs reproducible across machines. Bare mode doesn’t use your subscription login, so the job needs ANTHROPIC_API_KEY or an equivalent provider credential. Add --output-format json to get total_cost_usd per invocation, so cost per task becomes a number you track instead of a feeling. Full pipeline setup lives in headless Claude Code in CI and cron.

Recalibrate with the correction count

The ladder is not a chart you memorize; it is a loop you run. Pick a rung, run the task, count your corrections, and adjust the next pick.

The docs give the threshold: after two failed corrections on the same issue, stop correcting. The session’s context is polluted, and the initial calibration was wrong. Every correction after the second is correction tax on a decision you already know was mistaken.

Pro tip

After two failed corrections on the same issue, /clear and restart with a better initial prompt that folds in what you learned — usually one rung up from where you started. Rewriting the prompt beats arguing with a polluted context every time.

When a session goes badly, diagnose it with the docs’ three questions:

  • Context too noisy? Clear it, and consider a subagent firewall next time.
  • Prompt too vague? Stay on the same rung and write a better sentence.
  • Task too big for one pass? Climb one rung.

Each diagnosis maps to a different adjustment, which is what makes the ladder teachable rather than vibes. If you instrument your usage, correction counts and cost per rung become data, and your defaults drift toward calibrated instead of habitual.

The failure patterns you already know are miscalibrations by another name. The kitchen-sink session and infinite exploration are under-escalation — work that needed subagents or a /clear.

The custom-agent stack nobody needed is over-escalation. The ladder gives both a name and a fix.

On a team, calibration drifts person by person: one engineer plans everything, another fans out subagents for two-file reads, and nobody can compare notes because there is no shared vocabulary. When we roll Claude Code out across engineering teams, we adopt the ladder as that vocabulary, so “this is a rung-3 task” means the same thing in every review and retro. Making those defaults explicit is a core part of our Enable work.


Next steps

Related service

This is part of our Enable work — make the team able to operate it.