Claude Code PR automation: babysit to green

Set up Claude Code PR automation that takes pull requests from red to green: diagnose CI failures, answer review threads, and escalate on cue.

on this page

A PR that fails CI at 4 p.m. usually costs you the rest of the afternoon: fix, push, wait, repeat. This workflow replaces that with Claude Code PR automation.

By the end, you’ll have a Claude Code session that owns a PR — watching checks, pulling failed logs, fixing, pushing, and re-checking until it’s mergeable. You’ll also have an unattended GitHub Actions variant for when nobody is at the keyboard.

The runbook has three layers. An interactive babysit loop you run locally in a worktree. An unattended auto-fix workflow built on claude-code-action. And a judgment layer on top: an explicit policy for what Claude fixes and what gets escalated to you.

This walkthrough uses the direct Claude API and keeps the unattended workflow short. For the governed production version, with separate review and fix jobs, keyless model access and an audit trail, see Claude Code in CI/CD without handing it the keys.

What you’ll learn

  • Run a bounded babysit loop that takes a PR from red to green without you watching
  • Diagnose CI failures with gh pr checks, statusCheckRollup, and gh run view --log-failed
  • Handle review feedback across all three places GitHub stores it
  • Wire unattended auto-fix with workflow_run triggers, a narrow allowlist and layered loop prevention
  • Encode an auto-fix vs escalate policy so Claude stops instead of spinning

Prerequisites

  • Claude Code running daily in your workflow, with an authenticated GitHub CLI (gh auth status shows green)
  • A repository where you can add workflows and secrets
  • anthropics/claude-code-action@v1 set up via /install-github-app for the unattended half (setup is covered in the official docs, not here)
  • Working knowledge of GitHub Actions YAML and your repo’s branch protection rules

The interactive babysit loop

The default mode: you hand a PR to a local Claude Code session and stay available for escalations. Claude owns the mechanical loop.

Run the session in a dedicated worktree so the monitor-fix loop never collides with your active checkout. It’s the same isolation that makes parallel Claude Code sessions with git worktrees work — the babysitter gets its own working tree, and you keep coding in yours.

git worktree add ../pr-482 --detach
cd ../pr-482
gh pr checkout 482
claude

Then hand over the PR with a prompt that encodes the full loop, the wait discipline, and the stop conditions:

Babysit PR #482 until it's green and mergeable. Loop: check
`gh pr checks 482`, pull failed logs with
`gh run view <id> --log-failed`, fix, run the affected tests
locally, commit, push, then `gh pr checks 482 --watch` — never
a sleep loop. Also address unresolved review comments and reply
to each thread with what you changed. Stop and report to me if:
a failure predates this PR, a reviewer raises an architecture
question, or you've pushed 3 fixes and it's still red.
Do not merge.

Every clause in that prompt earns its place. Here is the loop it produces:

  1. 01

    Gather PR state in parallel

    Claude runs gh pr checks 482 for CI status, gh pr view 482 --json reviews,reviewDecision for approvals, and gh pr view 482 --json mergeStateStatus,mergeable for merge readiness. Unresolved inline threads come from the REST API. One pass, full picture.

  2. 02

    Triage in priority order

    CI failures first, review comments second, merge conflicts third. Red checks block everything else, so they get fixed before any review thread gets a reply.

  3. 03

    Pull failed logs and fix

    For each failing check, gh run view <run-id> --log-failed returns only the failing steps. Claude traces the failure to the diff and edits the code.

  4. 04

    Verify locally before pushing

    Run the affected tests locally — not the whole suite, the affected ones. A push that hasn’t been verified locally just burns a CI cycle. This is the core discipline from verification loops for agentic code: the fix isn’t done until it’s been observed passing.

  5. 05

    Commit new, never amend

    New commits, never --amend on pushed history. Reviewers need to see what changed since their last look, and amending destroys that trail.

  6. 06

    Watch, don't poll

    gh pr checks 482 --watch streams check status and exits on completion. When checks finish, loop back to step 1 — up to the iteration cap.

Pro tip

Say “watch, never poll” explicitly in the prompt or your skill. Left to improvise a wait, an agent can write a tight loop like until gh run view | grep -q completed; do true; done, and every iteration is a GitHub API call against your hourly rate limit. gh run watch <run-id> and gh pr checks <pr> --watch do the waiting for you instead of hand-rolled polling. In our experience, most babysit failures are self-inflicted: rate-limit lockouts and sleep loops that burn context.

For long-running CI, we use a stricter variant of the wait discipline: forbid the main agent from polling at all and push the CI watch — and the raw logs — into a background subagent. The parent session only receives a summary, which is the context firewall pattern applied to CI babysitting. Raw workflow logs are enormous; they don’t belong in the session that’s making judgment calls.

Danger

Force-push rules are non-negotiable and belong in every babysit prompt or skill. Never force-push to main or master, ever. On the PR branch, force-pushing is acceptable only for rebase-based conflict resolution, and only with --force-with-lease. And never let Claude pass --no-verify — pre-commit hooks are part of your verification, not an obstacle to it.

Diagnose CI failures from logs

The loop’s diagnostic engine is three gh commands. Prefer gh over the GitHub MCP server for this workflow — the official best practices guide calls CLI tools the most context-efficient way to interact with external services, and names gh for GitHub. Every tool an MCP server adds is context the loop carries, and on a loop that runs dozens of turns that line item belongs in your context budget.

Start with the check summary, then pull the failing logs:

gh pr checks 482
gh run view <run-id> --log-failed

gh pr checks lists each check with its status, duration, and a link to its run. gh run view --log-failed prints only the log lines from the failed steps, so a failing test arrives as its assertion and stack trace rather than the full job log. Say the test log shows an expected total that is off by one cent. The trail from there is short: the assertion points at a rounding change inside the PR’s own diff, Claude fixes it, runs that one test file locally, and pushes.

When check names don’t map cleanly to runs, enumerate them structurally:

gh pr view 482 --json statusCheckRollup \
  --jq '.statusCheckRollup[]'

Each entry carries a detailsUrl containing the run ID — extract it, then gh run view <run-id> --log-failed. This is more reliable than asking Claude to scrape run IDs from the human-readable gh pr checks table.

One diagnostic question comes before any fix: does this failure exist on main? If the same test fails on the base branch, no amount of fixing the PR will help, and pushing speculative fixes just muddies the history. That check is a stop condition, not a fix path.

Handle review feedback in all three places

Review feedback lives in three separate data sources, and a babysitter that reads only one will silently ignore reviewers. Gather all three:

# General PR conversation comments
gh pr view 482 --json comments

# Inline review comments (per-line threads)
gh api repos/acme/api/pulls/482/comments

# Review verdicts: approvals, change requests
gh pr view 482 --json reviews

Two rules make the feedback handling trustworthy rather than sycophantic. First, evaluate whether the feedback is actually valid before acting — reviewers are sometimes wrong, and a babysitter that blindly applies every suggestion is worse than none. On busy PRs, it can be worth spawning a stronger-model subagent per substantive comment just to make this call. Second, reply to every thread you act on, and sign the reply so humans can tell who wrote what.

Replying inside an inline thread requires the REST replies endpoint — gh pr comment can only post top-level comments:

gh api --method POST \
  "repos/acme/api/pulls/482/comments/2201389/replies" \
  -f body="Fixed in a3f81c2 — switched to banker's rounding as
suggested. 🤖 Written by Claude Code"

Substitute your owner, repo, PR number, and the comment ID from the inline-comments query above. This endpoint is why we codify review handling as explicit commands: natural-language instructions like “respond to the reviews” produce unpredictable results, while an explicit command sequence is repeatable.

Unattended Claude Code PR automation

The interactive loop assumes you’re reachable. The unattended half covers nights, weekends, and PRs you don’t own, using two GitHub Actions workflows built on claude-code-action@v1.

The first is the standard @claude mention responder. The action auto-detects interactive mode when someone mentions @claude in a PR comment or review thread, so reviewers can summon a fix on demand. The official docs cover that setup.

The second is the interesting one: auto-fix on CI failure, triggered by workflow_run. It runs only for failed CI on pull request branches in this repository, never for forks:

name: Auto Fix CI Failures
on:
  workflow_run:
    workflows: ["CI"]
    types: [completed]

permissions: {}   # nothing by default; the job grants its own

jobs:
  auto-fix:
    if: >-
      github.event.workflow_run.conclusion == 'failure' &&
      github.event.workflow_run.event == 'pull_request' &&
      github.event.workflow_run.head_repository.full_name == github.repository &&
      github.event.workflow_run.actor.type != 'Bot' &&
      !startsWith(github.event.workflow_run.head_branch, 'claude-auto-fix-ci-')
    runs-on: ubuntu-latest
    environment: claude-fix   # deployment branches: the default branch only
    timeout-minutes: 20       # placeholder: set from your own runs
    concurrency:
      group: claude-auto-fix-${{ github.event.workflow_run.head_branch }}
      cancel-in-progress: true
    permissions:
      contents: read
      actions: read
    steps:
      - name: Mint a GitHub App token so pushes re-trigger CI
        id: app-token
        uses: actions/create-github-app-token@v3
        with:
          client-id: ${{ vars.FIXER_APP_CLIENT_ID }}
          private-key: ${{ secrets.FIXER_APP_PRIVATE_KEY }}
          permission-contents: write

      # The default branch at the workspace root: the CLAUDE.md, settings,
      # hooks and .mcp.json that Claude Code loads come from here.
      - uses: actions/checkout@v7
        with:
          persist-credentials: false

      # The failing PR branch in pr/, pushed with the App token.
      - uses: actions/checkout@v7
        with:
          ref: ${{ github.event.workflow_run.head_branch }}
          path: pr
          token: ${{ steps.app-token.outputs.token }}

      - name: Commit as the App's bot user
        env:
          GH_TOKEN: ${{ steps.app-token.outputs.token }}
          APP_SLUG: ${{ steps.app-token.outputs.app-slug }}
        run: |
          BOT_ID=$(gh api "/users/${APP_SLUG}[bot]" --jq .id)
          git -C pr config user.name "${APP_SLUG}[bot]"
          git -C pr config user.email "${BOT_ID}+${APP_SLUG}[bot]@users.noreply.github.com"

      - name: Install dependencies without lifecycle scripts
        working-directory: pr
        run: npm ci --ignore-scripts

      # The failed job log, saved inside the workspace so Claude can read it
      # without web tools or gh.
      - name: Save the failed log inside the workspace
        env:
          GH_TOKEN: ${{ github.token }}
          RUN_ID: ${{ github.event.workflow_run.id }}
        run: |
          mkdir -p ci-log
          gh run view "$RUN_ID" --repo "$GITHUB_REPOSITORY" --log-failed > ci-log/failed.log

      - uses: anthropics/claude-code-action@v1
        with:
          # An environment secret on claude-fix, not a repository secret.
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          github_token: ${{ steps.app-token.outputs.token }}
          prompt: |
            CI failed on the branch checked out in ./pr. The failed job log
            is in ./ci-log/failed.log. Diagnose the failure and fix the
            root cause in ./pr.
            Run tests with: cd pr && npm test
            Run git only as: git -C pr diff, git -C pr add,
            git -C pr commit, git -C pr push
            If the failure looks flaky or unrelated to the change,
            change nothing and explain why.
          settings: |
            {
              "permissions": {
                "blockReadsOutsideWorkingDirectories": true,
                "allow": [
                  "Read(pr/**)",
                  "Read(ci-log/**)",
                  "Edit(pr/**)",
                  "Bash(npm test *)",
                  "Bash(npm run lint *)",
                  "Bash(git -C pr diff *)",
                  "Bash(git -C pr add *)",
                  "Bash(git -C pr commit *)",
                  "Bash(git -C pr push)"
                ],
                "deny": [
                  "Skill",
                  "Agent",
                  "Edit(.github/**)",
                  "Edit(pr/.github/**)",
                  "Bash(git * --force*)",
                  "Bash(git * -f)",
                  "Bash(git * -f *)",
                  "Bash(gh pr merge *)"
                ]
              }
            }
          claude_args: >-
            --permission-mode dontAsk
            --disallowedTools "WebFetch,WebSearch"
            --max-turns 10

This short variant leaves out protections the governed version adds, such as a --max-budget-usd cost cap alongside --max-turns; read Claude Code in CI/CD without handing it the keys before you run it unattended.

We keep this article on the direct Claude API. The key lives in ANTHROPIC_API_KEY as an environment secret on claude-fix, whose deployment branches are limited to the default branch, so a same-repo branch’s own pull_request workflow can’t read it the way it could read a repository secret. The App’s private key gets the same treatment. The stronger option is no stored key at all: the action supports workload identity federation for the Claude API through anthropic_federation_rule_id and anthropic_organization_id, which exchanges the job’s GitHub OIDC token instead and needs id-token: write on the job.

The layout is the part that’s easy to get wrong. We check out the default branch at the workspace root and the PR branch into pr/, because Claude Code loads CLAUDE.md, .claude/ settings, hooks, skills and .mcp.json from its working directory, and a branch can rewrite all of them. With the head at the root, the pull request would govern the agent that’s fixing it. Denying Skill and Agent covers the branch’s own skills and subagents, which Claude can still discover once it reads files under pr/, and treat the branch’s CLAUDE.md as untrusted input. The action’s security guide passes the PR checkout with --add-dir; we leave it out, because pr/ is already inside the working directory and --add-dir would load that branch’s skills at startup.

The allowlist is the other half. The old shape, Bash(npm:*) and Bash(git:*), isn’t a syntax problem: :* is just an older spelling of a trailing *, so Bash(npm:*) matches the same commands as Bash(npm *). The problem is scope. npm alone matches every npm command, including any script the branch adds to package.json, and git alone matches every subcommand, force push included. Putting the subcommand before the *, as in Bash(npm test *) or Bash(git -C pr commit *), limits the rule to that subcommand. The git rules name pr/ because a cd into another directory followed by git always prompts, since git there could run that directory’s hooks, and in dontAsk mode anything that would prompt is refused.

Deny rules always win over allow rules. Bash rules match the command text Claude writes, not every way of invoking a program, so sh -c variants slip past a deny rule; a short allow list under dontAsk is what actually closes the door. npm test still runs the branch’s own code, with the job’s credentials in reach, which is why this workflow runs only on branches pushed by people with write access to the repository. For same-repo branches, that write access is the real boundary.

The job’s permissions block stays minimal. actions: read is what lets the log step fetch the failed run’s log with the default token and save it to ci-log/failed.log. Without it, the agent is diagnosing blind: it has no web tools, no gh, and an App token that can only push. blockReadsOutsideWorkingDirectories keeps Claude’s file tools and built-in read commands inside the workspace, away from the push token actions/checkout stores under the runner’s temp directory; the Read rules document what it’s meant to read. Pushes go through the App token, so the default token needs nothing more than contents: read. Create a custom GitHub App with Contents, Issues and Pull requests read/write and no Workflows permission, and mint its token with contents write only.

Warning

The single most common reason unattended babysitting stalls: pushes made with the default GITHUB_TOKEN do not trigger downstream workflows. GitHub’s recursion guard means Claude’s fix lands but CI silently never re-runs, so the PR sits red forever. That is why the workflow above mints a token from your own GitHub App with actions/create-github-app-token and uses it for the PR checkout and the action. The official troubleshooting section says the same: remove a GITHUB_TOKEN passed to the action, or pass a custom app token instead.

Two trigger rules sit under all of this. Never combine pull_request_target with a checkout of the pull request’s code: it runs untrusted code with your secrets. And workflow_run runs the workflow file from the default branch, so a pull request can’t edit the fixer that acts on it.

Loop prevention is layered, because any one guard can fail:

  • Built-in bot rejection — the action refuses to run for bot actors unless they’re listed in allowed_bots, and we list none. It also requires write access from the actor who started the upstream CI run.
  • One attempt per human push — the fixer pushes with an App token, so its own commit re-runs CI as a bot. The actor.type != 'Bot' condition means a CI failure on Claude’s fix escalates to a person instead of triggering another fix, so the fixer can never chase its own output.
  • Branch-prefix guard — the !startsWith(head_branch, 'claude-auto-fix-ci-') condition covers the variant where fixes land on a separate claude-auto-fix-ci-* branch rather than the PR branch.
  • --max-turns 10 and timeout-minutes — a hard cap on agent turns per invocation, with the job timeout as the backstop, so a confused run can’t grind indefinitely.
  • A narrow allowlist under dontAsk — edits inside pr/, the test and lint commands, and four git subcommands. No force push, no edits under .github/, no web tools, no gh merge.
  • Concurrency control — the job’s concurrency group is keyed on the head branch with cancel-in-progress: true, so a newer failure cancels the older fixer instead of running two fixers on the same PR.

Note what this workflow deliberately doesn’t do: merge. Nothing in the allow list lets Claude merge or approve, and you shouldn’t route around that. Unattended Claude Code PR automation gets a PR to green; a human merges it, behind branch protection. Since these runs are headless Claude Code on a runner, the deeper flag-level coverage — output formats, permission modes, cost bounds — lives in headless Claude Code in CI and cron.

Auto-fix or escalate

Everything above is mechanism. This section is policy, and it’s what separates a babysitter from a menace. This is the line we draw:

SignalAction
Test failure with a clear assertionAuto-fix
Lint or format failureAuto-fix
Type errorAuto-fix
Build break traceable to the diffAuto-fix
Trivial merge conflictAuto-fix interactively (rebase, --force-with-lease, PR branch only); escalate when unattended
Flaky test needing infra contextEscalate
Environment-level failure (no code patch fixes it)Escalate
Architecture or design feedbackEscalate
Two reviewers who disagreeEscalate
Failure that predates the PREscalate
Still red after 3 iterationsEscalate

The iteration cap exists for a specific failure mode: Claude will spin hard on non-code problems, pushing speculative fix after speculative fix at a flaky test or a broken runner. Every one of those pushes is a correction you’ll pay for later — the correction tax compounds fastest when the agent can’t tell that the problem isn’t in the code. We cap it at 3 pushes, interactive or unattended: if three targeted fixes haven’t gone green, a fourth rarely will.

The stop conditions do real work in the prompt. Compare a naive babysit request with one that encodes the policy:

Before
Babysit PR #482 until CI is green. Fix whatever is failing and keep pushing until the checks pass.
After
Babysit PR #482 until CI is green. Fix, verify locally, push, watch. Stop and report if: the failure predates this PR, a reviewer raises an architecture question, the failure is infrastructure rather than code, or 3 pushes haven't gone green. Never merge, never force-push, never use --no-verify.

The before version has no exit besides success, which means its real exit is you noticing something’s wrong. The after version fails loudly and early. This matrix is one rung of a broader policy — where a task sits before you hand it off at all is covered in the escalation ladder for agentic tasks.

One governance gap to plan around: neither claude-code-action nor the CLI ships risk tiers. A dependency-bump fix and an auth-middleware rewrite look identical to the tooling. Encode that policy yourself — Edit path rules in the allow and deny lists, tighter caps on sensitive branches, and never auto-merging. And when the loop finishes, your review target shifts: you review the auto-fix commits, not the code that was broken. Reviewing AI diffs without rubber-stamping applies doubly to fixes an agent made to its own PR.

Package the loop as a skill

Prompting the loop from memory works once. Codifying your Claude Code PR automation as a skill makes it repeatable across sessions and teammates — the constraints ride along every time.

---
name: babysit-pr
description: Take a PR from red to green with a bounded fix loop.
---

# Babysit PR

Take PR $ARGUMENTS from failing to green and mergeable. Never merge.

## Gather state (run in parallel)

- `gh pr checks $ARGUMENTS`
- `gh pr view $ARGUMENTS --json reviews,reviewDecision`
- `gh pr view $ARGUMENTS --json mergeStateStatus,mergeable`
- `gh api repos/{owner}/{repo}/pulls/$ARGUMENTS/comments`
  for unresolved inline threads

## Triage order

1. CI failures
2. Review comments
3. Merge conflicts

## Loop rules

- Pull failed logs with `gh run view <run-id> --log-failed`.
- Reproduce and fix locally. Run the affected tests before pushing.
- New commits only. Never amend pushed history.
- Wait with `gh pr checks $ARGUMENTS --watch` or
  `gh run watch <run-id>`. Never write a polling or sleep loop.
- Maximum 3 iterations, then stop and report.

## Hard constraints

- Never force-push to main or master.
- Force-push to the PR branch only with --force-with-lease,
  only for rebase conflicts.
- Never skip pre-commit hooks (no --no-verify).
- Never merge. Report when green and mergeable.
- Reply to every review thread you act on, and sign replies.

## Escalate instead of fixing

- The same failure exists on main.
- A reviewer raises architecture or design questions.
- Reviewers disagree with each other.
- The failure is infrastructure or environment, not code.

The hard constraints in a skill are still just words in context — Claude follows them, but nothing enforces them. For the rules that must hold (--no-verify, force-pushing main), enforce them mechanically with a PreToolUse hook that blocks the command outright. That’s the difference described in hooks as guardrails, not suggestions, and a babysit loop that runs unattended is exactly where you want guardrails over suggestions.

The mechanism is portable; the matrix is not. Which failures an agent may fix on its own, how many pushes it gets, and which paths are off-limits depend on how your delivery process already handles review, CI ownership, and risk. Drawing that line for a specific pipeline is what we do when we assess where agents fit in a delivery process.


Next steps

Related service

This is part of our Assess work — decide what to build and why.