Claude Code prompt templates: spec-first

Copy three Claude Code prompt templates for bug fixes, features, and refactors, plus the four spec slots that let tasks land on the first pass.

on this page

What you’ll learn

  • Structure any task prompt around four slots: constraint, invariant, files in scope, and definition of done
  • Copy three Claude Code prompt templates tuned for bug fixes, features, and refactors
  • Generate a spec by interview when you cannot write one cold
  • Size a spec at 5-8 load-bearing lines, and see why longer prompts degrade
  • Reuse the spec as the review contract for the finished diff

Prerequisites

  • Claude Code in daily use on a real codebase. This guide skips setup and tool basics.
  • A test suite Claude can run. Every template here ends with a runnable check.
  • Working familiarity with plan mode, subagents, and hooks. Cross-links carry the depth.

Every unstated decision is a coin flip

“Implement a rate limiter for our API endpoints” reads like a complete request. It is the lazy prompt the official Claude Code best practices use in their Writer/Reviewer example, and it is actually a stack of unmade decisions: algorithm, scope, limit, storage backend, and whether a new dependency is fair game. Every pick is a coin flip.

The docs put it plainly: “Claude can infer intent, but it can’t read your mind.” The cost of the lazy prompt is deferred, not avoided; it arrives as the correction tax when you discover the in-memory token bucket you never asked for.

Spec-first prompting is spec-driven development shrunk to the size of one prompt. No toolkit, no constitution files. Four slots, and three templates your team can keep in the repo.

The four slots of a spec-first prompt

SlotWhat it carries
ConstraintDecisions you have already made
InvariantWhat must not change
Files in scopeWhere the work lives, and its boundary
Definition of doneA check Claude can run, plus evidence to show

Constraint is every decision you would otherwise correct later. The docs’ test-scoping example shows the shape: not “add tests for foo.py” but “write a test for foo.py covering the edge case where the user is logged out. avoid mocks.”

Invariant names what the diff must not touch. Anthropic’s help-center guide to giving Claude context uses compact forms like “keep the signature the same,” “do not edit generated/,” and “no new dependencies.” Invariants are cheap to write and expensive to omit.

Files in scope is where the problem lives and what Claude may touch: a boundary, not a line number. Pointing at a pattern beats describing one; the docs’ example is “HotDogWidget.php is a good example. follow the pattern.”

Definition of done is the highest-value line in the prompt. The docs are blunt about why: “Claude stops when the work looks done.” Give Claude a check (the test suite, the build exit code, a diff-against-fixture script) and it iterates until the check passes. Ask for evidence, not assertion: “show me the test output.” Stop hooks and verifier subagents are the escalation, covered in verification loops for agentic code.

One boundary matters: spec-first does not mean dictating the edit. The same help-center guide contrasts “Open userService.ts, find the validate function, add a null check on line 42” with “Users with no email are crashing the validation step. Make it handle that gracefully and add a test.” Specificity targets the four slots. The keystrokes stay Claude’s job.

Note

The four slots hold per-task decisions. Durable conventions — formatter, test command, commit style — belong in CLAUDE.md, where every session inherits them. Repeating them per prompt spends your instruction budget twice.

Claude Code prompt templates by task type

These three templates share the same four slots but weight them differently. Bug fixes lean on evidence, features on constraints, refactors on invariants. We keep them as files in a prompts/ directory so everyone starts from the same shape.

The bug-fix template

The official docs contrast the lazy and tight versions of the same bug report:

Before
fix the login bug
After
users report that login fails after session timeout. check the auth flow in src/auth/, especially token refresh. write a failing test that reproduces the issue, then fix it

Symptom, not diagnosis. Files in scope, with token refresh flagged. And a definition of done that forces reproduce-then-fix ordering: the failing test proves the bug existed and proves the fix works. Generalized:

[Symptom] What users see, not your theory of the cause.
[Evidence] Paste the full stack trace or failing output verbatim.
[Scope] The files or directories where you suspect the problem lives.
[Reproduce] Write a failing test that reproduces the issue, then fix it.
[Root cause] Address the root cause, don't suppress the error.
[Done] The new test and the existing suite pass. Show the output.

The root-cause clause comes from the docs’ own prompting table. Without it, “make the error go away” has a degenerate shortest path: swallow the exception.

Pro tip

Paste stack traces whole. Never summarize them. The help-center guide is explicit that the exact filename, line number, and message are what let Claude find the right location quickly.

The feature template

Back to the rate limiter, with all four slots filled:

Add rate limiting to the public API routes.

Constraint: sliding window, 100 requests/min per API key, backed by the
existing Redis client in src/lib/redis.ts. No new dependencies.

Invariant: authenticated internal routes (src/app/api/internal/**) are
untouched. Existing middleware signatures don't change.

Scope: src/middleware.ts, src/lib/rate-limit.ts (new), plus tests.
Nothing outside src/.

Done means: over-limit requests return 429 with a Retry-After header;
tests cover under-limit, over-limit, and window-reset; npm test and
npm run typecheck pass. Show me the test output.

Every decision the lazy prompt left to chance is now either made or fenced off, and Claude cannot call the work done until three named test cases pass. The template:

[Task] One sentence naming the outcome.
[Constraint] Decisions already made: algorithm, storage, limits,
  dependency policy.
[Invariant] What must not change: untouched routes, frozen signatures.
[Pattern] A file that already does it right: "follow src/lib/cache.ts."
[Scope] Files to create or modify, and the fence around them.
[Done] A check Claude can run, plus "show me the output."

The pattern slot deserves emphasis. One existing file that does it right outperforms paragraphs of style description.

The refactor template

Refactors are where the invariant slot carries all the weight. The task is, by definition, “change structure, preserve behavior”, so the spec’s job is to make “preserve behavior” checkable and to fence off opportunistic rewrites.

[Task] The structural change, in one sentence.
[Invariant] Behavior-preserving: existing tests pass without
  modification.
[Invariant] Public API frozen: exported signatures and return types
  unchanged.
[Scope] The module being restructured. Nothing outside it.
[No opportunism] Don't fix unrelated code, rename beyond the task, or
  "improve" tests to make them pass.
[Done] The untouched suite passes and the diff stays inside the named
  scope.

“Tests pass without modification” is the load-bearing line. A refactor that edits its own tests is grading its own homework. There is no official refactor template; this is our distillation of the constraint and invariant guidance into the refactor shape.

When you can’t write the spec cold

Sometimes you don’t know the constraints upfront. Make Claude extract them from you. This interview prompt is verbatim from the official docs:

I want to build [brief description]. Interview me in detail using the
AskUserQuestion tool.

Ask about technical implementation, UI/UX, edge cases, concerns, and
tradeoffs. Don't ask obvious questions, dig into the hard parts I might
not have considered.

Keep interviewing until we've covered everything, then write a complete
spec to SPEC.md.

Answer honestly, including “I don’t care”; that is a decision too. Then check SPEC.md against the four slots and cut anything that is neither a decision nor a check. The docs’ next step is explicit: “Once the spec is complete, start a fresh session to execute it.” The interview’s back-and-forth is spent context; see context budgeting.

For multi-file work inside a session, plan mode is the interactive version: the plan you approve is the spec you execute.

Keep the spec under eight lines

The failure mode on the other side is over-correcting from a lazy prompt to a 40-item checklist. Instruction-following research (the ManyIFEval and IFScale benchmarks) shows adherence degrades as instruction count rises: the chance of satisfying every instruction falls roughly as the per-instruction success rate raised to the power of the count.

Warning

Per-instruction reliability compounds. A spec-first prompt is 5-8 load-bearing lines, not a checklist. If a line isn’t a decision, an invariant, a boundary, or a check, cut it.

The line budget also decides what goes where. Durable conventions belong in CLAUDE.md, treated as a contract; the docs warn that “Bloated CLAUDE.md files cause Claude to ignore your actual instructions!” Invariants that must hold every single time, like “never edit generated/,” belong in hooks, which enforce rather than request. The prompt keeps only what is specific to this task.

And sometimes the right spec is no spec: “Vague prompts can be useful when you’re exploring and can afford to course-correct.” Matching spec effort to task size is its own skill; the escalation ladder covers it.

The spec is also the review contract

A spec pays twice: once when Claude executes it, and again when the diff comes back, because now there is something to review against. “Nothing outside the task’s scope changed” is only a checkable claim because the spec named the scope. Hand the spec and the diff to a fresh-context reviewer, as plan mode: get plans worth approving shows with a plan file, and the four slots become the review checklist.

Templates only pay off when they are shared. When we roll Claude Code out across engineering teams, a checked-in prompts/ directory like this one is among the first things we set up, because it turns one engineer’s good habit into the team default and gives reviewers a common standard to hold diffs against. Building that shared layer is what our Enable service is for.


Next steps

External resources:

Related service

This is part of our Enable work — make the team able to operate it.