Review AI-generated code in layers

Stop rubber-stamping giant diffs. Review AI-generated code with five layers: plans, tests, staged commits, adversarial sessions, and human eyes.

on this page

Claude Code finishes a session and hands you a 40-file diff. You scroll, the code looks plausible, and you click approve. That approval was a rubber stamp, and you know it.

Here is how we review AI-generated code honestly without reading every line, and without pretending we did. The answer is a layered strategy: review the plan, read the tests, checkpoint with commits, and send in a second Claude. Reserve your own eyes for the diffs that deserve them.

Skipping review does not make the cost disappear. It defers the correction tax to production, where it compounds.

What you’ll learn

  • Review implementation plans before code exists, then reuse them as conformance checklists
  • Read test diffs as the contract instead of auditing implementations line by line
  • Split one unreviewable session into staged commits sized under the human review ceiling
  • Run /code-review and adversarial subagent reviews in fresh context
  • Keep a short, blast-radius-keyed list of diffs that always get human eyes

Prerequisites

  • A Claude Code workflow that already uses plan mode, subagents, and hooks
  • A project with a real test suite Claude can run
  • Working git fluency: staging, ref ranges like main...feature, worktrees

You can’t review AI-generated code by volume

Human review has a measured ceiling. SmartBear’s well-known study of a Cisco team put effective review at roughly 400 changed lines; past that, defect detection drops sharply. Treat the number as an order of magnitude, not a law, but the direction is clear.

Agentic output blows past that ceiling routinely. As Bryan Finster argues in “AI Broke Your Code Review”, assistants multiply how much code a developer produces, while review capacity stays where it was. Anthropic’s Code Review launch post (March 2026) makes the risk concrete: on PRs over 1,000 changed lines, 84% of its automated reviews surfaced findings, averaging 7.5 issues per review.

Big diffs are where the bugs are. They are also exactly where human attention fails first.

The rubber stamp is cognitive, not moral. Reviewers approve code that looks syntactically right, even when they never evaluated its integration quirks or edge cases. The same post reports that before Anthropic deployed automated review internally, only 16% of its PRs received substantive review comments.

Neither of the obvious responses works. “Read faster” ignores the ceiling. “Trust the model” ignores the findings that pile up in large PRs.

The rest of this article treats your attention as a budget and spends it in layers, ordered by cost.

Layer 1: review the plan, not the diff

The cheapest review happens before any code exists. In plan mode, Claude reads your files and proposes an approach without editing anything. A 60-line plan is tractable in a way a 3,000-line diff never will be.

Plan review catches the most expensive failure class: solving the wrong problem. No amount of post-hoc diff reading recovers the hours lost to a correct implementation of a wrong design. If your plans are too vague to review, fix that first with plans worth approving.

Pro tip

Press Ctrl+G in plan mode to open the plan in your editor. Editing it directly is faster than negotiating changes in chat, and the edited plan becomes your conformance checklist for Layer 4.

The plan keeps paying after implementation. Scope drift, meaning files changed that the plan never mentioned, is the cheapest red flag to detect in a large diff. You will use the plan as a review artifact in Layer 4.

One caveat: planning has overhead. For a small, obvious change, skip the plan and most of this machinery too. Review rigor should scale with risk, which is Layer 5’s job to define.

Layer 2: read the tests as the contract

If you cannot read the implementation, read what the implementation claims about itself. Test names, assertions, and fixtures are a compressed statement of what the author believes the code does. If the belief is correct and the suite passes, the tests verify the implementation details mechanically.

This only works when the pass signal means something. Building trustworthy pass/fail signals is its own discipline, covered in verification loops for agentic code. This layer consumes those signals.

Warning

Tests written by the same session that wrote the code inherit the same misunderstanding. If Claude misread the requirement, the code and the tests are wrong together, and the suite passes anyway.

Three mitigations break that mirror:

  • Author a few acceptance cases yourself, from the requirement, before implementation starts. Spec-first prompting covers how to make those cases the source of truth.
  • Split the roles: have one Claude session write tests, then a different session write code to pass them. The official docs recommend this Writer/Reviewer variant directly.
  • Hunt for weakened and deleted tests. GitHub’s guidance names “tests that are deleted rather than fixed” as an AI-specific pitfall.

The weakened assertion is the artifact to grep for in every test diff:

Before
expect(res.status).toBe(403)
After
expect(res.status).toBeTruthy()

The second assertion passes for any 2xx, 4xx, or 5xx response. A diff that loosens assertions while the implementation changes is a diff hiding a behavior change.

Finally, demand evidence instead of claims. Have Claude show the test output and the commands it ran rather than asserting success. Reviewing evidence is faster than re-running the verification yourself.

Layer 3: staged commits as review checkpoints

A 40-file diff is usually four to six logical changes that happened to land at once. Make Claude commit at those boundaries and each commit becomes a review unit sized under the 400-line ceiling. Model, service layer, and route handler are three commits, not one.

Set the contract at plan time with this instruction:

Implement the plan one phase at a time. After each phase: run the test
suite, show me the output, and stop so I can review and commit before
you continue. Do not start phase 2 until I say "continue".

This converts one end-of-session mega-diff into a loop you control:

  1. 01

    Constrain the session up front

    Give Claude the phase-stop instruction above along with the approved plan. Each plan phase should end with “run tests, stop for review”.

  2. 02

    Review the phase in your git view

    Read the staged changes in your IDE’s git panel. One phase at a time, the diff fits your attention span. Make manual edits now if needed.

  3. 03

    Commit and hand back control

    Commit with a message naming the phase. Then tell Claude: “Committed as [hash]. Continue to phase 2.” The hash anchors the session to your checkpoint.

  4. 04

    Repeat until the plan is done

    Every commit stays under the review ceiling, and git bisect and rollback remain usable if something surfaces later.

Do not rely on memory for this discipline. Promote the rule into your project instructions, treating CLAUDE.md as a contract:

## Workflow

- Commit after each plan phase, with a message naming the phase.
- Run the test suite before each commit and show the output.
- Stop after each commit and wait for "continue" before proceeding.

If sessions still skip the rule under pressure, enforce it with a Stop hook instead of a sentence. That is the difference between guardrails and suggestions.

One distinction matters here. Claude Code’s /rewind checkpoints are automatic snapshots for in-session undo, and the docs are explicit that they are not a replacement for git.

Rewind checkpoints are recovery units. Commits are review units. You need both.

Layer 4: a second Claude as adversarial reviewer

Claude reviewing its own diff in-session shares the author’s mental model, the same failure as a human self-review. The official docs state the fix plainly: a fresh context improves code review because Claude won’t be biased toward code it just wrote.

A fresh subagent sees only the diff and the criteria you give it, never the reasoning that produced the change. That isolation is the entire point, and it is the same property that makes subagents work as context firewalls.

Check the diff against the plan

After Claude declares the work done, run a conformance review against the plan you approved in Layer 1:

Use a subagent to check the current diff against PLAN.md. For each plan
step, say whether it is implemented and where. List any edge case the plan
names that has no test, and any changed file the plan never mentions.
Skip style comments.

Those three checks — steps covered, edge cases tested, no out-of-scope changes — are exactly what you cannot see by skimming a 40-file diff. The subagent reports gaps back into the session, so Claude can fix and re-review without you ferrying findings between windows.

Run /code-review before opening the PR

The bundled /code-review command (alias /review) reviews a diff for correctness bugs in a background subagent with its own context, then returns findings to your session. By default it covers your branch’s commits ahead of upstream plus uncommitted changes. You can also pass a target: a file path, a PR number, a branch, or a ref range.

/code-review high main...feature/rate-limiter

Don’t confuse it with /simplify, which is a separate cleanup-only review that applies reuse and simplification fixes without hunting for bugs.

The leading argument is an effort level that tunes the trade-off: low and medium return fewer, higher-confidence findings, while high through max broaden coverage and may include uncertain ones. ultra hands the review to the deeper cloud ultrareview instead. Add --fix to apply findings to the working tree, or --comment to post them as inline PR comments.

Note

A reviewer prompted to find gaps will usually report some, even when the work is sound. Chasing every finding produces over-engineering: extra abstraction, defensive code, tests for impossible cases. Scope the reviewer to gaps that affect correctness or stated requirements.

Escalate to a full reviewer session

When the change is too large or risky for a one-shot subagent report, run the reviewer as its own session. A full session can be interrogated; a subagent report cannot. Run it from a separate git worktree so both sessions work cleanly in parallel.

Review the rate limiter implementation in @src/middleware/rateLimiter.ts.
Look for edge cases, race conditions, and consistency with our existing
middleware patterns.

Paste the findings back into the writer session with: “Here’s the review feedback: [Session B output]. Address these issues.”

For teams, the managed Code Review product runs a fleet of specialized reviewer agents on each PR, verifies candidate findings against actual code behavior, and posts severity-ranked comments inline. It reads a REVIEW.md file from the repo root and gives its contents to the agents that find and verify findings as your repository’s review instructions:

# Review calibration

- Flag only issues that affect correctness or stated requirements.
- Behavior claims need a file:line citation.
- Treat unscoped database queries, PII in logs or errors, and
  non-backward-compatible migrations as Important.
- Cap nit findings at three per review.

As of September 2026, the Code Review docs list it as a research preview for Team and Enterprise plans, billed on token usage; check that page for current availability and cost. Local /code-review does not read REVIEW.md, but it does follow your CLAUDE.md. Without the managed product, you get most of the value from /code-review plus the subagent patterns above. You can also move the same checks into CI with headless Claude Code and claude -p.

Layer 5: diffs that always get human eyes

Automated layers exist to concentrate human judgment, not replace it. Keep an explicit list of diff types where you still read every line, and key it on blast radius rather than line count:

  • Auth and security boundaries
  • Database migrations, especially anything not backward compatible
  • Code that moves money or touches billing
  • PII in logs, errors, or analytics
  • New dependencies
  • Anything irreversible: data deletion, external side effects, published APIs

Line count is a bad proxy for risk. Anthropic’s launch post describes Code Review flagging a single-line production change that would have disabled authentication for a service, the kind of change that looks routine in a diff. Small diff, maximum blast radius.

At the other end, routine formatting, docs, and comment-only diffs get the lightest tier: automation verifies, you approve. Everything in between gets layers matched to risk, the same logic as the escalation ladder applied to review.

Review AI-generated code with a budget

Put the layers together and you get a budget you can defend to your team. Every diff gets some layers. No diff needs all of them except the blast-radius list.

Diff typeLayers to apply
One-sentence fix, low riskTests pass, skim the diff
Feature in a planned sessionPlan review, test diff, staged commits
Large or unfamiliar surfaceAll of the above plus /code-review high
Multi-session or risky changeAdd plan-conformance subagent or reviewer session
Blast-radius list (auth, migrations, money, PII)Every layer, plus your eyes on every line

The point of the budget is honest approval. When you approve a diff under this system, you can say exactly which layers evaluated it and what each one checked.

That is a review. Scrolling and clicking approve never was.

When we help teams roll out Claude Code, this budget is usually the missing piece. Individual engineers adopt the tool fast, but nobody has agreed which diffs get which layers, so review quality depends on who happens to approve. Writing the blast-radius list and the layer table into shared config and team practice is a core part of our Enable work.


Next steps

External resources: Anthropic’s Claude Code documentation on code review and Bryan Finster’s “AI Broke Your Code Review” for the process-level diagnosis.

Related service

This is part of our Enable work — make the team able to operate it.