Claude Code stops when the work looks done, not when it is done. Without a check it can run, “looks done” is the only signal available — and every mistake waits for you to notice it.
A Claude Code verification loop closes that gap. You give the agent something that produces a pass or fail. Then you decide who enforces it: the model in the current turn, a committed test harness, a fresh-context reviewer, or a hook that refuses to end the turn.
We climb those four rungs here in order of increasing strength, with the trap waiting at each one.
What you’ll learn
- Recognize why agentic code fails “confidently wrong” and why looks-done carries zero signal
- Force evidence over assertion with prompt-level checks and failing-test-first repros
- Run TDD as a machine-readable harness, with safeguards against the agent editing tests
- Set up adversarial review in a fresh context and scope it to avoid over-engineering
- Gate turn endings deterministically with a Stop hook, and know when
/goalis enough
Prerequisites
- Claude Code v2.1.139 or later (
/goalrequires it) - Working knowledge of hooks and subagents; we cover only their verification use here
- A project with at least one real check: a test suite, a typechecker, or a linter
- Comfort editing
.claude/settings.jsonand writing short shell scripts
Why you need a Claude Code verification loop
An LLM generates code by predicting likely tokens. The output is syntactically valid, follows familiar patterns, and usually compiles. Then it fails on boundary conditions: empty strings, zero inputs, unhandled file I/O errors.
Research on coding-agent failures names the most dangerous category: code that is “syntactically valid, contextually plausible, and semantically wrong.” One taxonomy of LLM coding mistakes calls it near-miss syndrome.
This failure mode differs in kind from human bugs. A human who is unsure hedges, tests, or asks. An agent presents broken code with exactly the same confidence as working code, so “looks done” carries zero information.
Anthropic’s docs call the pattern the trust-then-verify gap, and the fix is blunt: always provide verification — tests, scripts, screenshots. If you can’t verify it, don’t ship it.
Until that verification exists, you are the loop. Every plausible-but-wrong “done” you catch by hand is part of the correction tax, and unlike the agent, you don’t scale.
The four-rung ladder
A check is anything that returns a signal Claude can read. A test suite, a build exit code, a linter, a script that diffs output against a fixture — all qualify. Once the check exists, the only question is who enforces it.
| Rung | Enforced by | Setup cost | Guarantee |
|---|---|---|---|
| 1. Verification in the prompt | Model attention | None | Advisory |
| 2. TDD harness | Committed tests | Minutes | Strong, tamper-visible |
| 3. Adversarial review | Fresh-context subagent | One prompt | Catches self-bias |
4. Stop hook / /goal | The harness itself | Script + settings | Deterministic |
Each rung trades setup cost for attention cost. Rung 1 costs nothing to set up and depends entirely on the model following instructions. Rung 4 needs a script and a settings entry, then fires whether or not anyone remembers it.
This ladder is a sibling of the escalation ladder for agentic tasks. That article escalates how much autonomy you grant; this one escalates verification strength to match.
Rung 1: verification in the prompt
The cheapest Claude Code verification loop is a sentence. Name the check, tell Claude to run it, and demand the output as evidence.
implement a validateEmail function
implement validateEmail. write failing tests first: empty string, missing @, unicode domains. run the tests, make them pass, and show me the output.
For bug fixes, use the reproduce-fix-prove pattern from the official docs:
users report that login fails after session timeout. check the auth flow in
src/auth/, especially token refresh. write a failing test that reproduces the
issue, then fix it. run the test suite and show me the output — address the
root cause, don't suppress the error.
Two phrases carry the weight. “Write a failing test that reproduces the issue” forces a red run before any fix, which proves the repro is real. “Show me the output” demands evidence instead of assertion.
Run this against a real bug and the transcript should show a red-then-green shape: the new test run once and failing before any fix, then the same test command passing after it. If you only ever see the green run, you have no proof the test reproduced the bug.
Insist on evidence even when you trust the work. Reading test output is faster than re-running the verification yourself, and it works for sessions you weren’t watching.
Rung 1’s limit: it is advisory. Nothing stops the model from skipping the run deep in a long, degraded context. When the check must always happen, keep climbing.
Rung 2: TDD as the agent’s harness
Anthropic’s post on building agents with the Claude Agent SDK calls rules-based feedback, such as linting, the best form of feedback, and describes LLM-as-judge as generally not very robust. Tests are the strongest form because they are a machine-readable spec. Each red-to-green cycle gives the agent unambiguous feedback it can iterate against without you in the loop.
Tests are the executable half of a spec. Spec-first prompting covers writing the prose half; this rung turns it into something the agent cannot argue with.
- 01
Write the tests first
Prompt Claude to write tests from the spec before any production code. Claude defaults to writing the code first, so be explicit: “write the tests, do not write the code yet.”
- 02
Confirm they fail
Run the suite and confirm red. A test that passes before the code exists proves nothing.
- 03
Commit the failing tests
This checkpoint is the safeguard the whole rung depends on. Agents sometimes edit tests to make them pass; with tests committed, tampering shows in the diff and reverts in one command.
- 04
Build to green without touching tests
Tell Claude to make the tests pass without modifying any test file. A permission deny rule or a hook guarding the test directory turns that request into a wall.
- 05
Diff the test files before merging
Run
git diff <checkpoint> -- tests/. Any change means the agent moved the goalposts instead of reaching them.
The official docs endorse a stronger split: one Claude session writes the tests, a second session writes the code to pass them. The test writer has never seen the code, so it can’t shape the tests around whatever got built.
Rung 3: adversarial review in a fresh context
A model reviewing code in the same conversation that produced it is grading its own homework. Its context is saturated with the reasoning that produced the code, so the same wrong assumptions carry straight into the review.
The fix is a reviewer subagent in a fresh context. It sees only the diff and the criteria you give it, not the reasoning behind the change — the same isolation that makes subagents work as context firewalls. The official docs put it as a tip: before treating a task as done, have a subagent review the diff in a fresh context and report gaps. In practice, name the work, the criteria, and what counts as a finding:
Before you call this done, have a subagent review the diff with fresh
context. Give it the task description above and the tests from rung 2.
It should report requirements that are unmet or untested. Report gaps,
not style preferences.
The subagent returns findings to the main session, which fixes and re-reviews without you relaying anything. If you worked from an approved plan, hand the reviewer that plan as its rubric, which is one more reason to get plans worth approving. For correctness-only review, the bundled /code-review skill packages this pattern, and takes an effort level (low through max) to trade coverage for confidence.
“Report gaps, not style preferences” is the load-bearing clause. A reviewer prompted to find gaps will usually report some, even when the work is sound. Chase every finding and you get over-engineering: extra abstraction layers, defensive code, tests for cases that can’t happen. Scope the reviewer to correctness and stated requirements; treat the rest as optional.
Practitioners push this further with cross-model review — a different model grades the diff, since same-model reviewers share training biases — and multi-persona panels. Treat both as extensions, not the baseline.
Keep one boundary straight: adversarial review shrinks your review burden but never removes it. Reviewing AI diffs without rubber-stamping covers the half of the loop that stays human.
Rung 4: gates that block the stop
Everything so far is advisory or model-judged. Hooks are the one verification mechanism the model cannot ignore. A CLAUDE.md line saying “run the tests” is a clause in a contract the model can forget; a Stop hook makes it non-optional. For general hook mechanics, see hooks as guardrails — this section covers only the verification gate.
A Stop hook fires when Claude tries to finish a turn. Register a script:
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/verify-turn.sh",
"timeout": 120
}
]
}
]
}
}
Then make the script run your checks and answer with an exit code:
#!/bin/bash
# Block the turn from ending until typecheck + scoped tests pass.
cd "$CLAUDE_PROJECT_DIR" || exit 0
if ! npm run typecheck --silent 2>/tmp/typecheck.log; then
echo "Typecheck failing — fix before finishing:" >&2
tail -20 /tmp/typecheck.log >&2
exit 2 # exit 2 blocks the stop; stderr goes back to Claude
fi
if ! npx vitest run --changed --silent 2>/tmp/tests.log; then
echo "Tests failing on changed files — fix before finishing:" >&2
tail -20 /tmp/tests.log >&2
exit 2
fi
exit 0 # checks clean — Claude may finish the turn
The exit code is the whole trick. Exit 0 allows the stop. Exit 2 blocks it, and everything printed to stderr feeds back to Claude as the reason — so Claude resumes work against the exact failure you showed it.
Exit code 1 does not block the stop. Only exit 2 — or JSON with decision set to "block" — gates the turn. A validator that exits 1 on failure looks like a gate in your settings and blocks nothing. Test the hook by forcing a failure and confirming the turn refuses to end.
Two design rules keep the gate sane. First, the check must be able to clear. An unconditional block would loop forever, so Claude Code caps it: after Stop hooks have continued the turn eight times in a row, it overrides the next block and ends the turn (the CLAUDE_CODE_STOP_HOOK_BLOCK_CAP environment variable raises the cap). Your script also receives stop_hook_active on stdin, set to true when Claude is already continuing because of a Stop hook, so it can avoid blocking on a condition that will never resolve. Second, keep it fast and scoped — changed files, one test file — because it runs at every turn end. The full suite belongs in CI.
Note that Vitest’s --changed flag assumes git state it can diff against. Check its behavior in your repo before trusting the gate.
PostToolUse hooks complement the Stop gate at finer grain. They can’t un-run a tool, but an exit 2 after an Edit feeds lint or typecheck failures back immediately instead of at turn end.
/goal: the model-judged gate
/goal sets a completion condition for the session. After each turn, a small fast model (Haiku by default) checks the condition against the transcript and returns yes or no with a reason. A “no” starts another turn with that reason as guidance.
/goal all tests in test/auth pass and lint is clean — prove it with `npm test`
and `npm run lint` output. Don't modify any test file. Stop after 20 turns.
Under the hood, /goal is a wrapper around a session-scoped, prompt-based Stop hook. The evaluator does not run commands or read files, so write conditions Claude’s own output can demonstrate. A good condition has one measurable end state, a stated check, a constraint, and a turn bound.
Choose between the two by scope and trust. /goal is session-scoped, typed ad hoc, and model-judged, so a persuasive transcript can fool it. A Stop hook is settings-scoped, permanent, and script-judged: use /goal for one-off conditions, a Stop hook for invariants.
Both run unattended. claude -p "/goal ..." drives the loop to completion in one invocation, which makes it the core building block for headless Claude Code in CI and cron.
Choosing your rung
Match the rung to the cost of a false “done.”
- Attended one-off task: rung 1. Ask for the check and the output; you’re watching anyway.
- Feature with a clear spec: rung 2. Committed failing tests are the spec’s teeth.
- Large diff you’ll be tempted to skim: add rung 3 before you read it, so the fresh context burns down the finding list first.
- Project invariants like “typecheck always passes”: rung 4 Stop hook. Invariants deserve determinism.
- Unattended or headless runs: rung 4 is mandatory. Nobody is present to be the loop.
Rungs stack. A strong default for serious work: a TDD harness for the task, a Stop hook holding typecheck plus scoped tests, and one adversarial review pass before your own. If you want data instead of instinct, instrument your Claude Code usage and count how often each gate actually fires.
The end state is a role change. You stop re-running the agent’s work and start reading its evidence: test output, commands with their exit codes, a reviewer’s findings. A Claude Code verification loop doesn’t make you trust the agent more — it makes trust unnecessary.
For a team, the rung-4 gates are where this pays off most. A Stop hook committed to the repo holds every engineer’s sessions to the same invariant, whether or not they have read this article. Deciding which checks become shared gates, and which stay personal habits, is a standard part of how we roll Claude Code out through Enable.
Next steps
- Babysit a PR to green with Claude Code — the end-to-end workflow that points these loops at a live CI pipeline
- Hooks as guardrails, not suggestions — full hook mechanics: events, matchers, JSON output, and settings scopes
- Reviewing AI diffs without rubber-stamping — the human half of the loop that no gate replaces
External references: the official Claude Code verification guidance and the hooks reference for exact Stop and PostToolUse semantics.