Incident Response Runbooks with Claude Code

Build structured incident response runbooks as Claude Code skills that check service health, gather logs, and suggest root causes.

on this page

What you’ll learn

  • Structure an incident response SKILL.md with severity levels, diagnostic steps, and escalation paths
  • Build health check modules that query service metrics through bash and MCP servers
  • Gather and correlate container logs with application metrics during an incident
  • Add safety guardrails using Claude Code hooks to prevent dangerous automated actions
  • Connect observability MCP servers (Datadog, Grafana) for real-time data access

Prerequisites

  • Claude Code installed and configured
  • Familiarity with building Claude Code skills (SKILL.md basics)
  • A running application stack with some form of monitoring (Prometheus, Datadog, or similar)
  • Basic understanding of incident severity levels and on-call workflows
  • Terminal access to your infrastructure (direct or via SSH)
  • jq installed (the safety hook uses it)

Your runbook wiki runs to dozens of pages. Nobody reads it during an incident. The on-call engineer gets paged at 2 AM, opens a terminal, and starts guessing which dashboard to check first.

Claude Code skills turn that static wiki into an interactive investigation agent. You define the diagnostic steps, severity classifications, and escalation paths in a SKILL.md file. When an incident hits, Claude follows the runbook — checking service health, pulling logs, correlating metrics, and presenting a root cause hypothesis with the evidence behind it.

This guide walks you through building an incident response skill from scratch, with the guardrails we insist on before one goes anywhere near production.

The problem with static runbooks

Traditional runbooks sit in Confluence or a Git repo as markdown files. They describe what to do during an outage, but they cannot do anything themselves. During a real incident, engineers face three problems:

Context switching kills speed. You read step 3 of the runbook, switch to Grafana, run a query, switch back to the runbook, read step 4, switch to the terminal, pull logs, switch back again.

Runbooks go stale. The service architecture changed six months ago. The runbook still references the old database host. Nobody updated it because nobody reads it until something breaks.

Correlation is manual. The runbook tells you to check the error rate, then check the database connections, then check recent deployments. But it cannot tell you that the error rate spiked exactly 30 seconds after the last deployment reduced the connection pool size.

A Claude Code skill addresses all three. It executes the diagnostic steps, pulls live data, and correlates findings in a single terminal session.

How we assess on-call readiness for agents

Before we’d let an agent near a pager rotation, we check four things. Keep them in mind as you build; each one maps to a section below.

  1. Investigation is read-only by construction. The credentials the agent runs with can read metrics, logs, and cluster state, and cannot change them. Prompt instructions are not a boundary.
  2. Destructive paths are blocked by something other than the skill text. A hook or a permission deny rule stops the dangerous command even if the model decides to run it.
  3. The runbook degrades gracefully. If the observability MCP server is down — which is likely during an incident — the skill falls back to scripts that still work.
  4. Every hypothesis cites evidence. The output names the metric, log line, or commit behind each claim, so the human on call can verify it in seconds.

Anatomy of an incident response skill

Create a directory for your skill and add the SKILL.md file:

mkdir -p .claude/skills/incident-response/scripts
touch .claude/skills/incident-response/SKILL.md

The skill needs three sections: metadata, investigation procedure, and escalation rules. Here is the full structure:

---
name: incident-response
description: |
  Structured incident response runbook. Use when investigating
  production incidents, service degradation, or alert escalation.
  Follows SEV1-SEV4 classification with diagnostic steps.
---

# Incident Response Runbook

## Severity classification

Before investigating, classify the incident:

| Level | Criteria | Response time | Escalation |
|-------|----------|---------------|------------|
| SEV1  | Full outage, data loss risk | Immediate | Page on-call lead + engineering manager |
| SEV2  | Major feature broken, >10% users affected | 15 minutes | Page on-call lead |
| SEV3  | Minor feature degraded, workaround exists | 1 hour | Slack notification |
| SEV4  | Cosmetic issue, no user impact | Next business day | Ticket only |

## Investigation procedure

Follow these steps in order. Do not skip steps.

### Step 1: Service health overview
Run health checks across all services. Report:
- HTTP error rates (5xx) per service
- P99 and P50 latency per service
- Database connection pool usage
- Memory and CPU utilization

### Step 2: Identify the blast radius
Determine which services are affected.
- Check upstream and downstream dependencies
- Identify whether the issue is isolated or cascading

### Step 3: Gather logs
Pull logs from affected services for the last 15 minutes.
- Look for exception patterns and stack traces
- Count error frequency and group by error type
- Check for connection timeout messages

### Step 4: Check recent changes
Look for deployments, config changes, or infrastructure
modifications in the last 2 hours.
- Git log for recent merges to main
- Deployment pipeline history
- Infrastructure change records

### Step 5: Correlate and diagnose
Connect the findings from steps 1-4:
- Did the error rate spike align with a deployment?
- Do the log errors point to a specific dependency?
- Is the issue resource exhaustion or a code defect?

Present a root cause hypothesis with supporting evidence.

### Step 6: Document findings
Summarize the incident with:
- Timeline of events
- Root cause (confirmed or suspected)
- Impact assessment
- Recommended remediation steps

## Escalation rules

- If SEV1: always escalate before attempting remediation
- If remediation requires config changes: get explicit approval first
- If root cause is unclear after 30 minutes: escalate to next tier
- Never restart production services without documenting the reason

This gives Claude a structured investigation path. Each step asks for specific, verifiable output rather than a vague summary.

Building the health check module

The skill above tells Claude what to check but not how to check it. You need to give Claude a way to gather real data. There are two approaches: bash scripts bundled with the skill, or MCP servers for your observability platform. We use both.

Bash scripts for direct access

Create a diagnostics script that Claude can run during investigation:

#!/bin/bash
# Service health check script
# Usage: ./health-check.sh [service-name]

SERVICE=${1:-"all"}
NAMESPACE=${KUBE_NAMESPACE:-"production"}

echo "=== Service Health Check ==="
echo "Timestamp: $(date -u +%Y-%m-%dT%H:%M:%SZ)"
echo "Target: $SERVICE"
echo ""

if [ "$SERVICE" = "all" ]; then
  echo "--- Pod Status ---"
  kubectl get pods -n "$NAMESPACE" \
    --no-headers \
    -o custom-columns=\
"NAME:.metadata.name,STATUS:.status.phase,\
RESTARTS:.status.containerStatuses[0].restartCount,\
AGE:.metadata.creationTimestamp"

  echo ""
  echo "--- Recent Events (warnings only) ---"
  kubectl get events -n "$NAMESPACE" \
    --field-selector type=Warning \
    --sort-by='.lastTimestamp' | tail -10
else
  echo "--- Pod Details: $SERVICE ---"
  kubectl get pods -n "$NAMESPACE" \
    -l "app=$SERVICE" -o wide

  echo ""
  echo "--- Resource Usage ---"
  kubectl top pods -n "$NAMESPACE" \
    -l "app=$SERVICE" 2>/dev/null || \
    echo "Metrics server not available"

  echo ""
  echo "--- Recent Logs (last 50 lines) ---"
  kubectl logs -n "$NAMESPACE" \
    -l "app=$SERVICE" --tail=50 \
    --since=15m 2>/dev/null || \
    echo "No logs available"
fi
Warning

This script requires kubectl configured with cluster access. For Docker Compose environments, replace kubectl commands with docker compose logs and docker compose ps. Match the script to your actual infrastructure, and run it with a read-only identity (for example, a Kubernetes service account bound to the built-in view role).

Add a second script for checking recent deployments:

#!/bin/bash
# Check recent deployments and changes
# Usage: ./recent-changes.sh [hours-back]

HOURS=${1:-2}
NAMESPACE=${KUBE_NAMESPACE:-"production"}

echo "=== Recent Changes (last ${HOURS}h) ==="
echo ""

echo "--- Newest ReplicaSets (one per rollout) ---"
kubectl get replicasets -n "$NAMESPACE" \
  --sort-by='.metadata.creationTimestamp' \
  -o custom-columns=\
"NAME:.metadata.name,CREATED:.metadata.creationTimestamp,DESIRED:.spec.replicas" \
  2>/dev/null | tail -10

echo ""
echo "--- Git Commits (last ${HOURS}h) ---"
git log --since="${HOURS} hours ago" \
  --format="%h %ai %s" 2>/dev/null || \
  echo "Not in a git repository"

echo ""
echo "--- Scaling Events ---"
kubectl get events -n "$NAMESPACE" \
  --field-selector reason=ScalingReplicaSet \
  --sort-by='.lastTimestamp' 2>/dev/null | tail -5

Make both scripts executable with chmod +x .claude/skills/incident-response/scripts/*.sh, then update your SKILL.md to reference them. Add this section after the investigation procedure:

## Available tools

When investigating, use these bundled scripts:

- `${CLAUDE_SKILL_DIR}/scripts/health-check.sh [service]` — Pod
  status, resource usage, recent logs, and warning events
- `${CLAUDE_SKILL_DIR}/scripts/recent-changes.sh [hours]` — Recent
  rollouts, git commits, and scaling events

Claude Code substitutes ${CLAUDE_SKILL_DIR} with the directory that holds the skill’s SKILL.md, so the paths work wherever the session starts.

MCP servers for observability platforms

For deeper metric access, connect Claude Code to your monitoring platform via MCP. Project-scoped MCP servers live in a .mcp.json file at the project root, which you commit so the whole on-call rotation gets the same setup. Don’t put them in .claude/settings.json; that file holds settings like permissions and hooks, not server definitions.

For Grafana, the official mcp-grafana server runs through uvx, Docker, or a release binary:

{
  "mcpServers": {
    "grafana": {
      "command": "uvx",
      "args": ["mcp-grafana", "--disable-write"],
      "env": {
        "GRAFANA_URL": "${GRAFANA_URL}",
        "GRAFANA_SERVICE_ACCOUNT_TOKEN": "${GRAFANA_SERVICE_ACCOUNT_TOKEN}"
      }
    }
  }
}
Tip

Keep --disable-write on during incident investigation. It removes the tools that change Grafana state — dashboard updates, alert rule changes, incident and annotation writes — while leaving the query tools in place.

Datadog ships a remote MCP server rather than a local package. Add it at project scope with the endpoint for your Datadog site from Datadog’s MCP setup docs; it authenticates with OAuth on first use:

claude mcp add --transport http --scope project datadog <your-datadog-mcp-endpoint>

That command writes the entry into .mcp.json for you. The Datadog server exposes write tools as well as read tools, so limit its toolsets to what the investigation needs and add deny rules for any write tools you don’t want Claude to call.

With an observability MCP server connected, update your SKILL.md investigation steps to use it:

### Step 1: Service health overview
Use the observability MCP server to query:
- Error rate: `sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)`
- Latency P99: `histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))`
- DB connections: `db_connections_active` vs `db_pool_size`

If no MCP server is available, run
`${CLAUDE_SKILL_DIR}/scripts/health-check.sh all` for a basic overview.

This dual approach — MCP servers as primary, bash scripts as fallback — keeps the runbook working when part of your observability stack is the thing that’s down.

Log gathering and correlation

The most valuable part of an AI-assisted runbook is correlation. A human engineer flips between logs and metrics in separate tools. Claude can hold both in one context and look for patterns across them.

Structure your SKILL.md so Claude gathers data before drawing conclusions:

## Correlation checklist

After gathering health metrics, logs, and recent changes,
answer these questions:

1. **Temporal correlation**: Did any metric change within
   5 minutes of a deployment or config change?
2. **Dependency chain**: Is the failing service dependent on
   another service that shows degraded metrics?
3. **Resource exhaustion**: Are connection pools, memory,
   or CPU at >90% utilization on any service?
4. **Error pattern**: Do log errors point to a single root
   cause (e.g., "connection refused", "timeout", "OOM")?
5. **Blast radius**: How many services show degraded metrics?
   One service suggests a local issue. Multiple services
   suggest infrastructure or shared dependency failure.

Present findings as:
- **Symptom**: What the user or alert reported
- **Evidence**: Specific metrics, log lines, and timestamps
- **Hypothesis**: Most likely root cause with confidence level
- **Next action**: Recommended fix or further investigation

Here is the shape of a run we’d expect, described rather than captured. You tell Claude the checkout service is returning 503s and ask it to run incident response. It classifies severity from the error rate, runs the health step and finds two degraded services that share a database, pulls logs showing connection-limit errors from both, and finds a deploy minutes before the spike that changed retry logic. The hypothesis it presents — connection exhaustion caused by that deploy — comes with the metric, the log lines, and the commit as evidence, and the recommended remediation (roll back, watch connection counts recover) stops for your approval.

Safety guardrails

An incident response skill that can execute commands needs strict boundaries. You do not want Claude restarting production databases at 3 AM because it seemed like a good idea.

Warning

Never give an incident response skill unrestricted write access to production infrastructure. Separate investigation (read-only) from remediation (requires approval). The skill should diagnose, not fix.

Pre-execution hooks

Use a PreToolUse hook to stop destructive commands before they run. Register it in .claude/settings.json:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          {
            "type": "command",
            "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/block-destructive.sh"
          }
        ]
      }
    ]
  }
}

Claude Code passes the pending tool call to the hook as JSON on stdin, with the shell command at .tool_input.command. Only exit code 2 blocks the call; any other non-zero exit is a non-blocking error and the command still runs. On exit 2, Claude Code feeds your stderr back to Claude as the reason:

#!/bin/bash
# PreToolUse hook: block destructive commands during incident response.
# Input arrives as JSON on stdin; the Bash command is at .tool_input.command.
command=$(jq -r '.tool_input.command // empty')

BLOCKED_PATTERNS=(
  'kubectl +delete'
  'kubectl +scale.*replicas=0'
  'docker +rm +-f'
  'docker +stop'
  'drop +table'
  'truncate'
  'rm +-rf'
)

for pattern in "${BLOCKED_PATTERNS[@]}"; do
  if grep -qiE -- "$pattern" <<<"$command"; then
    echo "Blocked: command matches destructive pattern '$pattern'." >&2
    echo "Incident response is read-only. Present the command for manual approval instead." >&2
    exit 2
  fi
done

exit 0

Make it executable with chmod +x .claude/hooks/block-destructive.sh. Pattern matching on command text is a tripwire, not a sandbox: a determined workaround such as a script that runs the same operation will get past it. That’s why we pair it with permission deny rules for the obvious commands (for example Bash(kubectl delete *)) and, more importantly, with read-only credentials, so the destructive call fails even if every other layer misses it. Hooks as guardrails covers the pattern in more depth.

Phased remediation

Structure your skill so investigation and remediation are separate phases. The investigation phase runs on its own. The remediation phase stops and waits for human approval.

Add this to your SKILL.md:

## Remediation protocol

After completing the investigation and presenting the root
cause hypothesis:

1. STOP and present the recommended remediation steps
2. Do NOT execute any remediation without explicit approval
3. If approved, execute one step at a time and verify each
4. After remediation, re-run Step 1 to confirm recovery

Remediation actions that ALWAYS require approval:
- Service restarts or redeployments
- Database operations (connection kills, failovers)
- Configuration changes to production services
- Scaling operations (up or down)
- Network changes (security groups, routing)
Pro tip

Run incident sessions in the default permission mode, where Claude asks before each shell command that isn’t already allowed, and add deny rules for the commands you never want run. Don’t rely on plan mode for this: plan mode is for researching and proposing changes, and approving a plan switches the session into a mode where Claude starts acting on it.

Connecting observability MCP servers

With your skill structure and safety guardrails in place, the final step is wiring everything together with your observability platform.

  1. 01

    Get the MCP server running

    For Grafana, install uv so uvx mcp-grafana can run, or switch the .mcp.json entry to the Docker image or a release binary.

    For Datadog there is nothing to install locally: the claude mcp add command above registers the remote server.

  2. 02

    Provide credentials through the environment

    The ${VAR} references in .mcp.json expand from the environment Claude Code starts in, so export the values in your shell (or a git-ignored .envrc loaded by direnv) before running claude:

    export GRAFANA_URL=https://your-instance.grafana.net
    export GRAFANA_SERVICE_ACCOUNT_TOKEN=your-grafana-service-account-token

    Use a Grafana service account with the Viewer role. Datadog authenticates through OAuth in the browser the first time Claude Code connects.

  3. 03

    Approve the project servers

    Start claude in the project. Claude Code asks you to approve project-scoped servers from .mcp.json before it uses them. Run claude mcp list to confirm they connect.

  4. 04

    Update your SKILL.md with MCP tool references

    Add a section to your skill that tells Claude which MCP tools to use for each investigation step. In our experience Claude follows a runbook more reliably when the skill maps each step to a named tool rather than leaving it to discover tools on its own. Tool names vary by server and version, so check what yours exposes with /mcp and substitute the real names for the placeholders below.

    ## Tool mapping
    
    | Investigation step | Primary tool | Fallback |
    |--------------------|-------------|----------|
    | Service health     | MCP metrics query tool | scripts/health-check.sh |
    | Error rates        | MCP metrics query tool | scripts/health-check.sh |
    | Container logs     | MCP log search tool    | kubectl logs |
    | Recent deploys     | MCP deployment/events tool | scripts/recent-changes.sh |
    | Traces             | MCP trace search tool  | (none) |
  5. 05

    Test the full workflow

    Run the skill against a staging incident or a recent resolved one before you trust it on a live page:

    $ claude
    > We're seeing elevated error rates on the API service.
      Run the incident response runbook.

    Claude should classify severity, run through each investigation step, and present a structured diagnosis. Break things on purpose too: stop the MCP server and confirm the skill falls back to the bash scripts, and try a blocked command to confirm the hook stops it.

Extending the skill

Once you have the base runbook working, consider these additions:

Service-specific runbooks. Create child skills for each critical service with service-specific PromQL queries, known failure modes, and historical incident patterns. Reference them from the parent skill: “If the affected service is payments, also load .claude/skills/incident-response-payments/SKILL.md.”

Post-mortem generation. Add a documentation step that produces a structured post-mortem from the investigation data. Claude already has the timeline, root cause, and impact assessment from the investigation — formatting it as a post-mortem template takes one additional instruction in the SKILL.md.

Integration with paging tools. Anthropic’s site reliability agent cookbook wraps PagerDuty and Confluence APIs as MCP tools for creating incidents and publishing post-mortems from an Agent SDK agent. The same approach works for any internal tool your on-call process depends on.

When we review an agent-assisted on-call setup, the runbook content is rarely the weak point. The gaps are in the layers around it: credentials broader than the investigation needs, guardrails that exist only as prompt text, and no fallback when the observability stack is itself degraded. Mapping those gaps against your real incident history is the kind of scoped review we run in an Assess engagement.


Next steps

Related service

This is part of our Assess work — decide what to build and why.