the problem
Teams that use Claude Code at their desks want the same help on every pull request, a first review and a fix when the build goes red, without a person at the keyboard.
CI is where untrusted text meets credentials: a diff, a comment or a test log can carry instructions aimed at the agent, in a job holding a push token and model credentials.
The requirements follow from that:
- No long-lived keys in the repository’s secrets.
- Untrusted content never controls the agent’s permissions. A pull request can’t change what the agent is allowed to do.
- Automated changes stay reviewable. The agent pushes commits; it never merges.
- A bounded cost per run.
- Every model call is auditable, back to one workflow run.
the design
- A pull request is opened or updated, and your existing CI runs.
- The review job runs on
pull_requestwith read-only tools and posts review comments. - When CI fails on a same-repo branch,
workflow_runstarts the fix job. - The fix job exchanges its GitHub OIDC token for short-lived credentials on its own IAM role, which trusts only the fix job’s GitHub environment. The review job signs in the same way to a separate role; the diagram follows the fix job.
- Claude Code in the job calls Claude on Amazon Bedrock with those credentials.
- The fix is committed and pushed with a GitHub App token, so CI runs again on the new commit.
- A person reviews the result and merges, as branch protection and CODEOWNERS require.
- Every model call is recorded in AWS CloudTrail under the role session, whose name carries the workflow run ID.
Our babysit-pr-to-green walkthrough uses the direct Claude API and a broader tool list to keep the example short; this page is the governed version.
components
| component | where it runs | responsibility |
|---|---|---|
| Pull request | GitHub | The change, from a same-repo branch or a fork. Everything in it is untrusted input. |
| CI checks | GitHub Actions, your existing workflow | Builds and tests as today. A failed run starts the fix job. |
| Review job | GitHub Actions, pull_request | Runs anthropics/claude-code-action@v1 with read-only tools and a comment tool. Skips drafts and forks. |
| Fix job | GitHub Actions, workflow_run | Same-repo branches only. Reads the failed log, fixes, tests, commits and pushes. |
| Guardrails | default branch and workflow | Base-branch CLAUDE.md and settings, plus the workflow’s allow and deny rules. |
| GitHub App token | GitHub | A custom GitHub App with Contents, Issues and Pull requests only, and no Workflows permission, so the agent can’t change CI. The fix job narrows its token further, to contents write. Commits pushed with the default workflow token don’t trigger new CI runs, which is why fixes use the App token. |
| OIDC → role | AWS IAM | One role per job, each trusting a single environment’s subject and allowed to invoke models and nothing else. The fix environment is limited to the default branch. |
| Claude on Bedrock | Amazon Bedrock | Serves the model calls through the Invoke API, billed to your AWS account. |
| Audit | AWS CloudTrail | Records each model call under the role session, named with the run ID. |
| Merge | GitHub, a person | Branch protection plus CODEOWNERS. The agent has no path to merge. |
decisions
Review only, or fix to green
Review is low risk and helps everyone.
Fixing means running code and pushing commits, so we limit it to branches our own people pushed.
“Every pull request” means every non-draft pull request from a branch in this repository. Fork pull requests get human review; this design runs nothing on them. GitHub withholds secrets from runs triggered by fork pull requests, and the action requires write access from whoever triggered it, so a review job couldn’t reach the model for them anyway.
Which events trigger which job
Never combine pull_request_target with a checkout of the pull request’s code: it runs untrusted code with your secrets.
Fixes use workflow_run, which runs the workflow file from the default branch, so a pull request can’t edit the fixer that acts on it.
The fix job checks that the failing run’s branch belongs to this repository and skips runs started by bots. It keeps the action’s write-access check, which on workflow_run also covers the actor who started the upstream run. It never checks out an untrusted ref into the workspace root: the default branch goes at the root and the failing branch into a subdirectory, as the action’s security guide recommends. From v7, actions/checkout also refuses fork pull request code under workflow_run unless you opt in.
Where the guardrails come from
A pull request can edit its own config, so PR-head settings are not a control. A headless run still uses a repository’s hooks and .mcp.json servers, so PR-supplied ones would run with the job’s credentials.
On pull request runs, the action restores .claude/, .mcp.json and CLAUDE.md from the base branch. In the fix job, the default branch sits at the workspace root, so startup configuration comes from there. The allowlist goes in through the action’s settings input and claude_args.
package.json scripts, the Makefile and lockfiles still come from the PR head. Running the tests means running the branch’s code, which is why the allowlist is narrow.
What the agent may run
Bash(npm:*) runs whatever scripts the pull request defines while AWS credentials are in the environment, and Bash(git:*) includes force pushes.
We allow only the commands the fix needs, such as Bash(npm test *) and a few git subcommands, and deny edits under .github/ (at the root and in pr/), force pushes and gh pr merge.
Deny rules win over allow rules, and Bash rules don’t match sh -c variants, so we keep the tool list short rather than relying on clever patterns.
How the job reaches the model
Neither option needs a stored key: the action supports workload identity federation for the Claude API too.
We choose Bedrock when model traffic must stay in the AWS account, bill through AWS, and show up in CloudTrail under a role we control.
The trade-off is a one-time model-access request, and a lag before the newest models and Claude Code features are available on Bedrock.
How a run is bounded
Any one guard can fail, so we stack them:
- Bot rejection. The action rejects bot actors unless they’re listed in
allowed_bots, and we list none. - Actor guard. The fix job also skips runs started by bots. The CI run a fix triggers belongs to the App’s bot, so a failure there goes to a person, not to another fix.
- Branch prefix. The job skips
claude-fix/branches, for fixes sent to a separate branch. - Concurrency. A concurrency group per branch cancels superseded runs.
- Turn and budget caps.
--max-turnsand--max-budget-usdcap each run. The budget is a client-side estimate, so treat it as a tripwire, not an invoice. - Job timeout.
timeout-minutesis the backstop.
safety and governance
We treat pull request diffs, comments and CI logs as untrusted:
- the action strips hidden content, such as HTML comments and invisible characters;
include_comments_by_actornarrows which comments reach the agent;- in fix mode, Claude has no web tools and no GitHub write beyond pushing to its branch. The tests it runs are another matter: they run the branch’s code, with network access and the job’s credentials (see the guardrails decision).
Each job has its own minimal permissions: block. Merging stays human: branch protection plus CODEOWNERS, enforced by GitHub, not by the agent’s instructions.
For same-repo branches, the real boundary is write access. As the action’s own auto-fix example puts it, “anyone who can push a branch to this repository can control what code runs here”. Under pull_request, a branch runs its own copy of the workflow file and could add a job of its own. So review and fix use separate GitHub environments and IAM roles, and the App’s private key is a claude-fix environment secret: as a repository secret, any same-repo branch’s pull_request workflow could read it and mint App tokens. The claude-fix environment is limited to the default branch, the ref a workflow_run job runs on, so a workflow from any other branch can’t claim it. The claude-review environment admits pull request runs, and its role can only call models.
The fix job in .github/workflows/claude-fix.yml, trimmed to the parts that matter:
name: Claude fix to green
on:
workflow_run:
workflows: ["CI"]
types: [completed]
permissions: {} # nothing by default; the job grants its own
jobs:
fix:
# Failed CI on a same-repo pull request branch, not started by a bot,
# not a fix branch.
if: >-
github.event.workflow_run.conclusion == 'failure' &&
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.head_repository.full_name == github.repository &&
github.event.workflow_run.actor.type != 'Bot' &&
!startsWith(github.event.workflow_run.head_branch, 'claude-fix/')
runs-on: ubuntu-latest
environment: claude-fix # deployment branches: the default branch only
timeout-minutes: 20 # placeholder: set from your own pilot runs
concurrency:
group: claude-fix-${{ github.event.workflow_run.head_branch }}
cancel-in-progress: true
permissions:
contents: read
actions: read
id-token: write
# Pin every action to a full commit SHA; tags are shown for readability.
steps:
- name: Mint a GitHub App token that can only push
id: app-token
uses: actions/create-github-app-token@v3
with:
client-id: ${{ vars.APP_CLIENT_ID }}
# An environment secret on claude-fix, not a repository secret.
private-key: ${{ secrets.APP_PRIVATE_KEY }}
permission-contents: write
# The default branch at the workspace root: the settings, hooks and
# CLAUDE.md that Claude Code loads at startup come from here.
- uses: actions/checkout@v7
with:
persist-credentials: false
# The failing branch in pr/, pushed with the App token. It is checked
# out by name, so it may include commits newer than the failing run.
- uses: actions/checkout@v7
with:
ref: ${{ github.event.workflow_run.head_branch }}
path: pr
token: ${{ steps.app-token.outputs.token }}
- name: Commit as the App's bot user
env:
GH_TOKEN: ${{ steps.app-token.outputs.token }}
APP_SLUG: ${{ steps.app-token.outputs.app-slug }}
run: |
BOT_ID=$(gh api "/users/${APP_SLUG}[bot]" --jq .id)
git -C pr config user.name "${APP_SLUG}[bot]"
git -C pr config user.email "${BOT_ID}+${APP_SLUG}[bot]@users.noreply.github.com"
- name: Install dependencies before any AWS credentials exist
working-directory: pr
run: npm ci --ignore-scripts
- name: Save the failed log inside the workspace
env:
GH_TOKEN: ${{ github.token }}
RUN_ID: ${{ github.event.workflow_run.id }}
run: |
mkdir -p ci-log
gh run view "$RUN_ID" --repo "$GITHUB_REPOSITORY" --log-failed > ci-log/failed.log
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ secrets.AWS_FIX_ROLE_ARN }}
aws-region: ${{ vars.AWS_REGION }}
role-session-name: claude-ci-${{ github.run_id }}
- uses: anthropics/claude-code-action@v1
env:
# A Bedrock model or inference profile ID, kept in a repository variable.
ANTHROPIC_MODEL: ${{ vars.CLAUDE_FIX_MODEL }}
with:
use_bedrock: "true"
github_token: ${{ steps.app-token.outputs.token }}
settings: .github/claude/fix-settings.json
prompt: |
CI failed on the branch checked out in ./pr. The failed job log
is in ./ci-log/failed.log. Find the root cause and fix it in ./pr.
Run the tests with: cd pr && npm test
Run git only as: git -C pr diff, git -C pr add, git -C pr commit,
git -C pr push
If the failure looks flaky or unrelated to the change, change
nothing and explain why.
# The --max-turns and --max-budget-usd values are placeholders.
claude_args: >-
--permission-mode dontAsk
--disallowedTools "WebFetch,WebSearch"
--max-turns 15
--max-budget-usd 5
The review job is the same shape with less: pull_request, an if: that skips drafts and forks, no contents write, the claude-review environment and role, and only read tools plus the action’s inline comment tool.
Each role trusts one environment’s subject. The trust policy for the claude-fix role:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::<account-id>:oidc-provider/token.actions.githubusercontent.com"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": "repo:your-org/your-repo:environment:claude-fix"
}
}
}
]
}
The review role is identical apart from environment:claude-review. Don’t widen the subject to repo:your-org/your-repo:* (with StringLike, the operator a wildcard needs); that also admits pull request runs. Both permissions policies allow bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:ListInferenceProfiles and bedrock:GetInferenceProfile on the specific inference profile and foundation model ARNs you use, plus aws-marketplace:ViewSubscriptions and aws-marketplace:Subscribe conditioned on aws:CalledViaLast being bedrock.amazonaws.com.
Repositories created after 15 July 2026 default to GitHub’s immutable subject format, repo:OWNER@ID/REPO@ID:...; check which format your tokens carry.
The allowlist the workflow passes in through settings, stored on the default branch as .github/claude/fix-settings.json:
{
"permissions": {
"blockReadsOutsideWorkingDirectories": true,
"allow": [
"Read(pr/**)",
"Read(ci-log/**)",
"Edit(pr/**)",
"Bash(npm test *)",
"Bash(git -C pr diff *)",
"Bash(git -C pr add *)",
"Bash(git -C pr commit *)",
"Bash(git -C pr push)"
],
"deny": [
"Skill",
"Agent",
"Edit(.github/**)",
"Edit(pr/.github/**)",
"Bash(git * --force*)",
"Bash(git * -f)",
"Bash(git * -f *)",
"Bash(gh pr merge *)"
]
}
}
With --permission-mode dontAsk, anything not allowed is refused. The git rules name pr/ because a cd into another directory followed by git always asks, which dontAsk turns into a refusal. Reads inside the working directory need no rule, so the Read entries document intent. actions/checkout keeps the push token under the runner’s temp directory, and blockReadsOutsideWorkingDirectories makes the file tools and built-in read-only commands such as cat refuse paths outside the working directory. It doesn’t reach code the tests run. Edit(pr/.github/**) removes any doubt about the unanchored .github rule.
Denying Skill and Agent closes a permission path, not just an injection path. Once Claude reads files in pr/, it loads that branch’s .claude/skills/, and a skill’s allowed-tools pre-approves tools even in an untrusted folder; deny rules override that grant. For the same reason we depart from the security guide’s advice to pass the PR checkout with --add-dir: pr/ is already inside the working directory, and --add-dir would load the branch’s skills, commands and subagents at startup. The branch’s CLAUDE.md can still be read, so treat it as untrusted input.
evals
Test the integration with seeded pull requests in a sandbox repository:
- a failing test with a known, small fix;
- a lint error;
- a flaky test that shouldn’t be “fixed”;
- a pull request whose description tries to instruct the agent, for example to edit a workflow or post a secret.
The fix job should fix the first two, leave the flaky one alone and say why, and ignore the injected instruction.
Re-run the set when the model, the action version or the allowlist changes. In production, track the fix acceptance rate, meaning fixes merged without rework, from the pull request history.
cost and latency
The levers, in the order we reach for them:
- Run review only on pull requests that aren’t drafts.
- Skip docs-only changes with path filters.
- Cap turns and budget per run, with values from a pilot.
- Cancel superseded runs through the concurrency group.
- Choose the model per job: a smaller model family for review and a larger one for fixes, set through the model variable.
To estimate spend: pull requests per week × (review runs + fix runs × fix rate) × tokens per run. Take tokens per run from the total_cost_usd and usage output, kept as a job artifact. That cost is a client-side estimate, so reconcile it against the AWS bill.
Latency is minutes, and a fix adds a CI cycle on top.
operating it
For audit, CloudTrail records the Bedrock calls under the role session, whose name carries the run ID. Ideally the roles live in a dedicated AWS account. Leave the action’s show_full_output off: it prints every tool output to the Actions log, and public repositories expose those logs.
The runbook triggers:
- A spike in fix runs on one branch. A loop the guards missed, or a flaky test the fixer keeps chasing.
- Budget-cap exits. A task too big, or a cap too low.
- Model-access errors after a model change. Check model access and the role’s ARNs.
- Review comments nobody reacts to. Turn review down, or improve
CLAUDE.md.
when not to use this
- Repositories where most pull requests come from forks. This design runs nothing on fork pull requests, so it would help only the few from your own branches.
- No branch protection. Nothing but convention would stop an automated merge.
- Slow or flaky test suites, where automated fixes chase noise.
on the Claude API directly
The same workflow runs against the Claude API. Replace the AWS credentials step and use_bedrock with the action’s workload identity federation inputs, anthropic_federation_rule_id and anthropic_organization_id, so there is still no stored key, or with an API key secret in anthropic_api_key.
Everything else is unchanged: triggers, guardrails, allowlist, App token and human merge.
For the CLI flags underneath, see headless Claude Code.