You already review diffs before merging. Claude Code plan mode gives you an earlier, cheaper gate: at the approval prompt, approach and scope are still free to change.
Most users rubber-stamp plans the way they rubber-stamp diffs. This guide teaches you to make a plan earn its yes.
What you’ll learn
- Judge a plan against the six elements every approvable plan contains
- Run a six-step interrogation checklist before pressing approve
- Iterate on the plan with annotation prompts instead of fixing code later
- Pick the approval option that matches how hard you vetted the plan
- Decide when plan mode is overhead and skip it without guilt
Prerequisites
- Claude Code CLI, mid-2026 release or later
- Fluency with the basics:
Shift+Tabmode cycling, permission modes, and a CLAUDE.md that works as a contract - A real multi-file change to practice on — this is not a toy-example guide
Plan approval is a review, not a formality
Every agentic session has exactly one moment where the approach is fully negotiable and nothing has been built: the plan approval prompt. After approval, every correction costs more — you are reverting edits, re-running tests, and untangling half-applied changes. That escalating cost is the correction tax, and plan review is the cheapest point on the curve.
Yet plan review gets less scrutiny than diff review. A wall of confident markdown appears, it looks thorough, and you approve. The problem: a plan can be plan-shaped filler — fluent, structured, and wrong about your codebase.
Treat the plan the way you treat a design doc from a new teammate. Interrogate its assumptions before any of them become code. Reviewing the diff is your second gate; the plan is your first, and it is the only gate where “do it completely differently” costs nothing.
What approving a plan actually does
Approval is not just a “yes” — it is a mode switch. In plan mode, Claude reads files and runs exploratory commands but cannot edit anything. Start a session in it directly:
claude --permission-mode plan
Or make it the default for a project where unreviewed edits are unacceptable:
{
"permissions": {
"defaultMode": "plan"
}
}
When Claude presents a finished plan, it asks how to proceed. Per the permission modes docs, the choices are:
- Yes, and use auto mode: approve and execute in auto mode. Where auto mode isn’t available to your session, this option auto-accepts edits instead.
- Yes, manually approve edits: approve and review each edit individually.
- No, keep planning: stay in plan mode and tell Claude what to change.
Exact labels shift between versions, so check your CLI. Each approve option sets the permission mode for execution. That makes approval a trust decision, not a binary.
The rule: how hard you interrogated the plan determines which option you pick. A plan that survived three rounds of scrutiny can run in auto mode. A plan you skimmed deserves manual approval of each edit, or “keep planning.”
Press Ctrl+G at the approval prompt to open the plan in your default editor. Editing the plan directly is faster than describing changes in chat, and Claude proceeds from your edited version.
One honesty note. Armin Ronacher’s teardown of plan mode argues it is largely prompt scaffolding plus a markdown plan file the agent maintains. The permission system independently blocks edits, but there is no magic in the mode. The value is the reviewable artifact and the checkpoint, which means the discipline transfers to any plan-file workflow.
Set "showClearContextOnPlanAccept": true in your settings and the approval list gains a first option that approves the plan and clears the planning context. It is off by default. Take it for large plans: executing on a fresh context beats dragging exploration debris along, for the reasons covered in context budgeting.
Anatomy of a plan worth approving
A plan worth approving names six things. If any are missing, the plan is not done — send it back.
- Change surface. Every file it expects to touch, including tests, callers, and config.
- Per-file logic. What gets added or removed in each file, with code snippets for non-obvious parts.
- Side effects. Migrations, new environment variables, new dependencies.
- Risks and trade-offs. What could break, and what it chose not to do.
- Success criteria. How you and Claude will know it worked.
- Phased todo list. Ordered tasks you can track during execution.
The difference between filler and substance shows up at the bullet level:
Update the auth code to support backup codes
Modify src/auth/session.ts to persist backup codes in Redis; adds migration 0042; new env var REDIS_URL
The first bullet tells you nothing you could verify. The second names a file, a storage decision, a migration, and a config change — four claims you can check against reality before approving.
Success criteria deserve special attention because you can’t count on a plan to include them. They come from you, the same way a spec does. Plans and specs are siblings: the spec constrains what you ask for, the plan reveals what Claude intends to do about it — demand both.
Interrogate the plan before you approve
Run this checklist on every plan before touching the approval menu. It takes two to five minutes and targets the specific ways plans go wrong.
- 01
Verify it read the real files
Ask: “Which files did you read to produce this plan?” A plan built from pattern-memory instead of your codebase will hallucinate helpers, paths, and conventions. If the answer is vague, send it back to explore first.
- 02
Audit the change surface
Does the file list include tests, callers, and config — or just the obvious file? A plan that touches
session.tsbut notsession.test.tshas not thought about the change, only the code. - 03
Hunt for duplication
Ask: “Does anything in this plan duplicate logic that already exists?” The classic failure is an endpoint or utility that re-creates something one directory away. Plans that ignore existing caching layers or shared validators pass in isolation and break the system.
- 04
Test its assumptions about state
Every plan assumes things about storage, schema, and auth. Pick the two assumptions the plan depends on most and verify them yourself. This step catches the expensive failures — see the example below.
- 05
Check the scope against your ask
Is the plan doing what you asked, or what it decided you meant? Refactors, “while we’re here” cleanups, and dependency bumps you never requested get cut now — not discovered in the diff.
- 06
Demand a definition of done
Ask: “Add explicit success criteria to the plan: what passes, what stays unchanged, how we verify. Don’t implement yet.” If Claude cannot state what done means, it cannot verify its own work later.
One caution from Anthropic’s docs, aimed at review subagents, is equally true of you. A reviewer prompted to find gaps will usually report some, even when the work is sound. Interrogate for correctness and scope — do not generate churn to feel thorough.
Example: two assumptions caught before code
Here is the shape of what step 4 catches. Say the task is adding TOTP-based MFA to an internal tool, and the plan comes back looking thorough. Two assumptions are worth checking by hand.
First, the plan persists MFA backup codes through session storage. If that storage is in-memory, backup codes vanish on every restart. Second, the plan assumes the user model already has an mfa_enabled field. If it doesn’t, the build hits an improvised migration halfway through.
Both problems are architectural, not syntactic. No linter, test, or diff review flags them cleanly; they surface as a half-broken deployment. Caught at the plan, each is a one-line correction plus an added success criterion: existing tests pass, new tests cover the TOTP and backup-code paths, and the non-MFA login flow is unchanged.
That is the trade plan interrogation offers: minutes of reading against hours of untangling. The plan is where wrong assumptions are still just sentences.
Iterate on the plan, not the code
When interrogation finds problems, do not approve-and-patch — fix the plan. Boris Tane’s annotation cycle is the sharpest public version of this discipline. In paraphrase: Claude writes a detailed plan.md grounded in the real source files; you annotate it inline in your editor; Claude addresses every note and updates the document; repeat until the plan survives review; then ask for a phased todo list. Every one of those prompts ends with the same suffix: “don’t implement yet.”
The annotations are where the value lands. A note like “this field belongs on the parent record, not on each item” is a data-model error corrected in a markdown file, before any migration exists. Each cycle costs one prompt; the same error caught in a finished diff costs a revert, a schema change, and a re-run of everything downstream.
The “don’t implement yet” suffix is load-bearing. Omit it and Claude starts coding the moment it judges the plan good enough — which defeats the entire checkpoint. Append it to every plan-iteration prompt.
Inside built-in plan mode, the equivalent loop is “Keep planning” plus written feedback, or Ctrl+G for direct edits. Tane prefers a plan file on disk for full control; the built-in mode gives you the permission gate for free. They are two versions of the same discipline — pick by ergonomics, not ideology.
When Claude Code plan mode is overhead
Planning is not free. A plan-annotate-approve cycle on a five-minute task turns it into a twenty-five-minute task, and no amount of caught assumptions justifies that on a typo fix. Anthropic’s docs draw the line explicitly: planning pays when you are uncertain about the approach, when the change spans multiple files, or when the code is unfamiliar.
The official heuristic is worth memorizing: “If you could describe the diff in one sentence, skip the plan.” Typo fixes, log lines, and renames go straight to execution.
Two more cases where you should skip it. Exploratory spikes, where you want Claude to just try something and you plan to throw the result away. And well-worn changes in code you know cold, where your own review of the diff is fast and reliable.
The cost model: planning tokens are cheap and wrong code is expensive, but your attention is the scarcest resource in the loop. Plan mode is one rung on an oversight ladder that runs from direct prompt to plan to full spec — the escalation ladder covers how to pick the rung. Match the rung to the blast radius of being wrong.
The plan outlives approval
A plan you interrogated hard keeps paying after you approve it. Three downstream uses make the effort compound.
First, it becomes a verification contract. Once the code exists, hand the plan to a fresh-context reviewer — the pattern from Anthropic’s docs, verbatim:
Use a subagent to review the rate limiter diff against PLAN.md. Check that every requirement is implemented, the listed edge cases have tests, and nothing outside the task’s scope changed. Report gaps, not style preferences.
The plan you interrogated becomes the rubric a subagent grades the diff against. This closes the loop described in verification loops, and a fresh-context subagent grades it without the coding session’s bias toward its own work.
Second, the plan is a lightweight decision record. It captures why the change took this shape — cheaper than a formal ADR and better than archaeology through commit messages.
Third, it is compaction insurance. The context window docs list the plan Claude wrote in plan mode among the things re-injected from disk after compaction, so it re-anchors the model after summarization drops conversational nuance — one of the recovery anchors covered in surviving auto-compact.
One last habit: when plan review catches a wrong assumption, write the correction into CLAUDE.md. “Session storage is in-memory; nothing durable goes there” costs one line and saves every future session from replanning the same mistake. That is CLAUDE.md working as a contract, fed by your plan reviews.
Across a team, plan review is the habit that decays first: it is invisible in the diff, so nobody notices when engineers stop doing it. When we roll Claude Code out across engineering teams, we make it visible instead, with a shared interrogation checklist, an agreed rule for which approval option fits which level of scrutiny, and plan corrections flowing back into the shared CLAUDE.md. That is the kind of groundwork our Enable service covers.
Next steps
- The correction tax in agentic coding — the cost model that makes plan review the highest-value gate
- Verification loops for agentic code — turning your approved plan into the rubric that grades the diff
- The escalation ladder for agentic tasks — choosing between direct prompts, plan mode, and full specs
External resources: the official plan mode docs, Boris Tane’s How I use Claude Code for the full annotation cycle, and Armin Ronacher’s What is plan mode? for the under-the-hood teardown.