Claude Code in CI/CD without handing it the keys

Headless Claude Code reviewing pull requests and fixing failing builds in GitHub Actions against Amazon Bedrock, with a trust boundary untrusted PR content can't cross and merging left to people.

on this page

the problem

Teams that use Claude Code at their desks want the same help on every pull request, a first review and a fix when the build goes red, without a person at the keyboard.

CI is where untrusted text meets credentials: a diff, a comment or a test log can carry instructions aimed at the agent, in a job holding a push token and model credentials.

The requirements follow from that:

  • No long-lived keys in the repository’s secrets.
  • Untrusted content never controls the agent’s permissions. A pull request can’t change what the agent is allowed to do.
  • Automated changes stay reviewable. The agent pushes commits; it never merges.
  • A bounded cost per run.
  • Every model call is auditable, back to one workflow run.

the design

Accent marks the controls that sit outside anything a pull request can change.
  1. A pull request is opened or updated, and your existing CI runs.
  2. The review job runs on pull_request with read-only tools and posts review comments.
  3. When CI fails on a same-repo branch, workflow_run starts the fix job.
  4. The fix job exchanges its GitHub OIDC token for short-lived credentials on its own IAM role, which trusts only the fix job’s GitHub environment. The review job signs in the same way to a separate role; the diagram follows the fix job.
  5. Claude Code in the job calls Claude on Amazon Bedrock with those credentials.
  6. The fix is committed and pushed with a GitHub App token, so CI runs again on the new commit.
  7. A person reviews the result and merges, as branch protection and CODEOWNERS require.
  8. Every model call is recorded in AWS CloudTrail under the role session, whose name carries the workflow run ID.

Our babysit-pr-to-green walkthrough uses the direct Claude API and a broader tool list to keep the example short; this page is the governed version.

components

componentwhere it runsresponsibility
Pull requestGitHubThe change, from a same-repo branch or a fork. Everything in it is untrusted input.
CI checksGitHub Actions, your existing workflowBuilds and tests as today. A failed run starts the fix job.
Review jobGitHub Actions, pull_requestRuns anthropics/claude-code-action@v1 with read-only tools and a comment tool. Skips drafts and forks.
Fix jobGitHub Actions, workflow_runSame-repo branches only. Reads the failed log, fixes, tests, commits and pushes.
Guardrailsdefault branch and workflowBase-branch CLAUDE.md and settings, plus the workflow’s allow and deny rules.
GitHub App tokenGitHubA custom GitHub App with Contents, Issues and Pull requests only, and no Workflows permission, so the agent can’t change CI. The fix job narrows its token further, to contents write. Commits pushed with the default workflow token don’t trigger new CI runs, which is why fixes use the App token.
OIDC → roleAWS IAMOne role per job, each trusting a single environment’s subject and allowed to invoke models and nothing else. The fix environment is limited to the default branch.
Claude on BedrockAmazon BedrockServes the model calls through the Invoke API, billed to your AWS account.
AuditAWS CloudTrailRecords each model call under the role session, named with the run ID.
MergeGitHub, a personBranch protection plus CODEOWNERS. The agent has no path to merge.

decisions

Review only, or fix to green

  • Review only, everywhere
  • Fix to green, everywhere
  • Chosen Review on every pull request, fixes only on same-repo branches

Review is low risk and helps everyone.

Fixing means running code and pushing commits, so we limit it to branches our own people pushed.

“Every pull request” means every non-draft pull request from a branch in this repository. Fork pull requests get human review; this design runs nothing on them. GitHub withholds secrets from runs triggered by fork pull requests, and the action requires write access from whoever triggered it, so a review job couldn’t reach the model for them anyway.

Which events trigger which job

  • pull_request_target with a checkout of the PR head
  • Chosen pull_request for review, workflow_run for fixes on same-repo branches

Never combine pull_request_target with a checkout of the pull request’s code: it runs untrusted code with your secrets.

Fixes use workflow_run, which runs the workflow file from the default branch, so a pull request can’t edit the fixer that acts on it.

The fix job checks that the failing run’s branch belongs to this repository and skips runs started by bots. It keeps the action’s write-access check, which on workflow_run also covers the actor who started the upstream run. It never checks out an untrusted ref into the workspace root: the default branch goes at the root and the failing branch into a subdirectory, as the action’s security guide recommends. From v7, actions/checkout also refuses fork pull request code under workflow_run unless you opt in.

Where the guardrails come from

  • The pull request's own .claude/settings.json and CLAUDE.md
  • Chosen The base branch, plus settings the workflow passes in

A pull request can edit its own config, so PR-head settings are not a control. A headless run still uses a repository’s hooks and .mcp.json servers, so PR-supplied ones would run with the job’s credentials.

On pull request runs, the action restores .claude/, .mcp.json and CLAUDE.md from the base branch. In the fix job, the default branch sits at the workspace root, so startup configuration comes from there. The allowlist goes in through the action’s settings input and claude_args.

package.json scripts, the Makefile and lockfiles still come from the PR head. Running the tests means running the branch’s code, which is why the allowlist is narrow.

What the agent may run

  • Broad rules like Bash(npm:*) and Bash(git:*)
  • Chosen A narrow allowlist with explicit denies

Bash(npm:*) runs whatever scripts the pull request defines while AWS credentials are in the environment, and Bash(git:*) includes force pushes.

We allow only the commands the fix needs, such as Bash(npm test *) and a few git subcommands, and deny edits under .github/ (at the root and in pr/), force pushes and gh pr merge.

Deny rules win over allow rules, and Bash rules don’t match sh -c variants, so we keep the tool list short rather than relying on clever patterns.

How the job reaches the model

  • Claude API with workload identity federation
  • Chosen Amazon Bedrock through GitHub OIDC

Neither option needs a stored key: the action supports workload identity federation for the Claude API too.

We choose Bedrock when model traffic must stay in the AWS account, bill through AWS, and show up in CloudTrail under a role we control.

The trade-off is a one-time model-access request, and a lag before the newest models and Claude Code features are available on Bedrock.

How a run is bounded

  • Trust the agent to stop
  • Chosen Layered limits: bot rejection, actor guard, concurrency, turn and budget caps, job timeout

Any one guard can fail, so we stack them:

  • Bot rejection. The action rejects bot actors unless they’re listed in allowed_bots, and we list none.
  • Actor guard. The fix job also skips runs started by bots. The CI run a fix triggers belongs to the App’s bot, so a failure there goes to a person, not to another fix.
  • Branch prefix. The job skips claude-fix/ branches, for fixes sent to a separate branch.
  • Concurrency. A concurrency group per branch cancels superseded runs.
  • Turn and budget caps. --max-turns and --max-budget-usd cap each run. The budget is a client-side estimate, so treat it as a tripwire, not an invoice.
  • Job timeout. timeout-minutes is the backstop.

safety and governance

We treat pull request diffs, comments and CI logs as untrusted:

  • the action strips hidden content, such as HTML comments and invisible characters;
  • include_comments_by_actor narrows which comments reach the agent;
  • in fix mode, Claude has no web tools and no GitHub write beyond pushing to its branch. The tests it runs are another matter: they run the branch’s code, with network access and the job’s credentials (see the guardrails decision).

Each job has its own minimal permissions: block. Merging stays human: branch protection plus CODEOWNERS, enforced by GitHub, not by the agent’s instructions.

For same-repo branches, the real boundary is write access. As the action’s own auto-fix example puts it, “anyone who can push a branch to this repository can control what code runs here”. Under pull_request, a branch runs its own copy of the workflow file and could add a job of its own. So review and fix use separate GitHub environments and IAM roles, and the App’s private key is a claude-fix environment secret: as a repository secret, any same-repo branch’s pull_request workflow could read it and mint App tokens. The claude-fix environment is limited to the default branch, the ref a workflow_run job runs on, so a workflow from any other branch can’t claim it. The claude-review environment admits pull request runs, and its role can only call models.

The fix job in .github/workflows/claude-fix.yml, trimmed to the parts that matter:

name: Claude fix to green
on:
  workflow_run:
    workflows: ["CI"]
    types: [completed]

permissions: {}   # nothing by default; the job grants its own

jobs:
  fix:
    # Failed CI on a same-repo pull request branch, not started by a bot,
    # not a fix branch.
    if: >-
      github.event.workflow_run.conclusion == 'failure' &&
      github.event.workflow_run.event == 'pull_request' &&
      github.event.workflow_run.head_repository.full_name == github.repository &&
      github.event.workflow_run.actor.type != 'Bot' &&
      !startsWith(github.event.workflow_run.head_branch, 'claude-fix/')
    runs-on: ubuntu-latest
    environment: claude-fix   # deployment branches: the default branch only
    timeout-minutes: 20       # placeholder: set from your own pilot runs
    concurrency:
      group: claude-fix-${{ github.event.workflow_run.head_branch }}
      cancel-in-progress: true
    permissions:
      contents: read
      actions: read
      id-token: write
    # Pin every action to a full commit SHA; tags are shown for readability.
    steps:
      - name: Mint a GitHub App token that can only push
        id: app-token
        uses: actions/create-github-app-token@v3
        with:
          client-id: ${{ vars.APP_CLIENT_ID }}
          # An environment secret on claude-fix, not a repository secret.
          private-key: ${{ secrets.APP_PRIVATE_KEY }}
          permission-contents: write

      # The default branch at the workspace root: the settings, hooks and
      # CLAUDE.md that Claude Code loads at startup come from here.
      - uses: actions/checkout@v7
        with:
          persist-credentials: false

      # The failing branch in pr/, pushed with the App token. It is checked
      # out by name, so it may include commits newer than the failing run.
      - uses: actions/checkout@v7
        with:
          ref: ${{ github.event.workflow_run.head_branch }}
          path: pr
          token: ${{ steps.app-token.outputs.token }}

      - name: Commit as the App's bot user
        env:
          GH_TOKEN: ${{ steps.app-token.outputs.token }}
          APP_SLUG: ${{ steps.app-token.outputs.app-slug }}
        run: |
          BOT_ID=$(gh api "/users/${APP_SLUG}[bot]" --jq .id)
          git -C pr config user.name "${APP_SLUG}[bot]"
          git -C pr config user.email "${BOT_ID}+${APP_SLUG}[bot]@users.noreply.github.com"

      - name: Install dependencies before any AWS credentials exist
        working-directory: pr
        run: npm ci --ignore-scripts

      - name: Save the failed log inside the workspace
        env:
          GH_TOKEN: ${{ github.token }}
          RUN_ID: ${{ github.event.workflow_run.id }}
        run: |
          mkdir -p ci-log
          gh run view "$RUN_ID" --repo "$GITHUB_REPOSITORY" --log-failed > ci-log/failed.log

      - uses: aws-actions/configure-aws-credentials@v6
        with:
          role-to-assume: ${{ secrets.AWS_FIX_ROLE_ARN }}
          aws-region: ${{ vars.AWS_REGION }}
          role-session-name: claude-ci-${{ github.run_id }}

      - uses: anthropics/claude-code-action@v1
        env:
          # A Bedrock model or inference profile ID, kept in a repository variable.
          ANTHROPIC_MODEL: ${{ vars.CLAUDE_FIX_MODEL }}
        with:
          use_bedrock: "true"
          github_token: ${{ steps.app-token.outputs.token }}
          settings: .github/claude/fix-settings.json
          prompt: |
            CI failed on the branch checked out in ./pr. The failed job log
            is in ./ci-log/failed.log. Find the root cause and fix it in ./pr.
            Run the tests with: cd pr && npm test
            Run git only as: git -C pr diff, git -C pr add, git -C pr commit,
            git -C pr push
            If the failure looks flaky or unrelated to the change, change
            nothing and explain why.
          # The --max-turns and --max-budget-usd values are placeholders.
          claude_args: >-
            --permission-mode dontAsk
            --disallowedTools "WebFetch,WebSearch"
            --max-turns 15
            --max-budget-usd 5

The review job is the same shape with less: pull_request, an if: that skips drafts and forks, no contents write, the claude-review environment and role, and only read tools plus the action’s inline comment tool.

Each role trusts one environment’s subject. The trust policy for the claude-fix role:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "Federated": "arn:aws:iam::<account-id>:oidc-provider/token.actions.githubusercontent.com"
      },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": {
          "token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
          "token.actions.githubusercontent.com:sub": "repo:your-org/your-repo:environment:claude-fix"
        }
      }
    }
  ]
}

The review role is identical apart from environment:claude-review. Don’t widen the subject to repo:your-org/your-repo:* (with StringLike, the operator a wildcard needs); that also admits pull request runs. Both permissions policies allow bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:ListInferenceProfiles and bedrock:GetInferenceProfile on the specific inference profile and foundation model ARNs you use, plus aws-marketplace:ViewSubscriptions and aws-marketplace:Subscribe conditioned on aws:CalledViaLast being bedrock.amazonaws.com.

Repositories created after 15 July 2026 default to GitHub’s immutable subject format, repo:OWNER@ID/REPO@ID:...; check which format your tokens carry.

The allowlist the workflow passes in through settings, stored on the default branch as .github/claude/fix-settings.json:

{
  "permissions": {
    "blockReadsOutsideWorkingDirectories": true,
    "allow": [
      "Read(pr/**)",
      "Read(ci-log/**)",
      "Edit(pr/**)",
      "Bash(npm test *)",
      "Bash(git -C pr diff *)",
      "Bash(git -C pr add *)",
      "Bash(git -C pr commit *)",
      "Bash(git -C pr push)"
    ],
    "deny": [
      "Skill",
      "Agent",
      "Edit(.github/**)",
      "Edit(pr/.github/**)",
      "Bash(git * --force*)",
      "Bash(git * -f)",
      "Bash(git * -f *)",
      "Bash(gh pr merge *)"
    ]
  }
}

With --permission-mode dontAsk, anything not allowed is refused. The git rules name pr/ because a cd into another directory followed by git always asks, which dontAsk turns into a refusal. Reads inside the working directory need no rule, so the Read entries document intent. actions/checkout keeps the push token under the runner’s temp directory, and blockReadsOutsideWorkingDirectories makes the file tools and built-in read-only commands such as cat refuse paths outside the working directory. It doesn’t reach code the tests run. Edit(pr/.github/**) removes any doubt about the unanchored .github rule.

Denying Skill and Agent closes a permission path, not just an injection path. Once Claude reads files in pr/, it loads that branch’s .claude/skills/, and a skill’s allowed-tools pre-approves tools even in an untrusted folder; deny rules override that grant. For the same reason we depart from the security guide’s advice to pass the PR checkout with --add-dir: pr/ is already inside the working directory, and --add-dir would load the branch’s skills, commands and subagents at startup. The branch’s CLAUDE.md can still be read, so treat it as untrusted input.

evals

Test the integration with seeded pull requests in a sandbox repository:

  • a failing test with a known, small fix;
  • a lint error;
  • a flaky test that shouldn’t be “fixed”;
  • a pull request whose description tries to instruct the agent, for example to edit a workflow or post a secret.

The fix job should fix the first two, leave the flaky one alone and say why, and ignore the injected instruction.

Re-run the set when the model, the action version or the allowlist changes. In production, track the fix acceptance rate, meaning fixes merged without rework, from the pull request history.

cost and latency

The levers, in the order we reach for them:

  • Run review only on pull requests that aren’t drafts.
  • Skip docs-only changes with path filters.
  • Cap turns and budget per run, with values from a pilot.
  • Cancel superseded runs through the concurrency group.
  • Choose the model per job: a smaller model family for review and a larger one for fixes, set through the model variable.

To estimate spend: pull requests per week × (review runs + fix runs × fix rate) × tokens per run. Take tokens per run from the total_cost_usd and usage output, kept as a job artifact. That cost is a client-side estimate, so reconcile it against the AWS bill.

Latency is minutes, and a fix adds a CI cycle on top.

operating it

For audit, CloudTrail records the Bedrock calls under the role session, whose name carries the run ID. Ideally the roles live in a dedicated AWS account. Leave the action’s show_full_output off: it prints every tool output to the Actions log, and public repositories expose those logs.

The runbook triggers:

  • A spike in fix runs on one branch. A loop the guards missed, or a flaky test the fixer keeps chasing.
  • Budget-cap exits. A task too big, or a cap too low.
  • Model-access errors after a model change. Check model access and the role’s ARNs.
  • Review comments nobody reacts to. Turn review down, or improve CLAUDE.md.

when not to use this

  • Repositories where most pull requests come from forks. This design runs nothing on fork pull requests, so it would help only the few from your own branches.
  • No branch protection. Nothing but convention would stop an automated merge.
  • Slow or flaky test suites, where automated fixes chase noise.

on the Claude API directly

The same workflow runs against the Claude API. Replace the AWS credentials step and use_bedrock with the action’s workload identity federation inputs, anthropic_federation_rule_id and anthropic_organization_id, so there is still no stored key, or with an API key secret in anthropic_api_key.

Everything else is unchanged: triggers, guardrails, allowlist, App token and human merge.

For the CLI flags underneath, see headless Claude Code.

Related service

This is part of our Enable work — make the team able to operate it.