Claude Code headless mode in CI and cron

Run Claude Code headless in CI pipelines and cron jobs with JSON output, deny-by-default permissions, and bounded cost. Includes three full recipes.

on this page

Every interactive Claude Code session has the same bottleneck: a person at the keyboard. Claude Code headless mode removes it. claude -p runs the full agent loop, prints the result, and exits — schedulable, pipeable, and CI-shaped.

This guide treats agentic coding as infrastructure, which is how we run these jobs ourselves: every unattended Claude Code job gets the same checks as any other service we put on a scheduler. You will build unattended jobs that are stateless, least-privileged, bounded in cost, and observable. Then you will apply the pattern in three complete recipes.

What you’ll learn

  • Run claude -p with --bare for deterministic, context-free CI invocations
  • Parse the JSON result envelope and enforce typed output with --json-schema
  • Lock down unattended runs with --permission-mode dontAsk and scoped allow rules
  • Bound every job with --max-turns, wall-clock timeouts, and lock files
  • Ship three recipes: CI failure triage, nightly dependency PRs, and release notes

Prerequisites

  • A current Claude Code CLI v2.1.x release: v2.1.205 or later for strict --json-schema validation, and v2.1.259 or later if you use --permission-prompts
  • Daily experience with interactive Claude Code sessions
  • An ANTHROPIC_API_KEY (or Bedrock/Vertex credentials) available to your runner
  • GitHub CLI (gh) and GitHub Actions for the recipes (the commands port to GitLab, Jenkins, or crontab)
  • Working knowledge of jq and shell scripting

Claude Code headless mode is an API

The official docs now frame print mode as the Agent SDK via the CLI. Pass -p (or --print) with a prompt and Claude runs the complete agent loop: reading files, executing tools, iterating. It prints the final result to stdout and exits.

Most CLI options combine with -p — the exceptions, such as --bg, fail with an error naming the conflict. The ones you will use constantly are --allowedTools, --output-format, --permission-mode, --append-system-prompt, --max-turns, and --continue/--resume for session chaining.

The flag that makes headless runs reproducible is --bare:

claude --bare -p "Summarize this file" --allowedTools "Read"

Bare mode skips auto-discovery of hooks, skills, plugins, MCP servers, auto memory, and CLAUDE.md. Without it, claude -p loads the same context an interactive session would — including whatever sits in a teammate’s ~/.claude or the project’s .mcp.json. The docs call bare mode the recommended mode for scripted calls and say it will become the default for -p in a future release.

This is a determinism argument. A CI job should produce the same behavior on every runner, and --bare starts from zero: Bash, file read, and file edit tools only. Everything else is opt-in via --append-system-prompt-file, --settings, --mcp-config, --agents, or --plugin-dir.

Think of --bare as context budgeting taken to its limit — you add only what the job needs. It also sidesteps a hidden tax: plain -p loads CLAUDE.md on every invocation, so a bloated project memory file bills you on every cron tick. That is one more reason to keep CLAUDE.md a tight contract and pass job-specific instructions explicitly.

One auth note: bare mode skips OAuth and keychain reads. Credentials must come from ANTHROPIC_API_KEY or your Bedrock/Vertex/Foundry provider configuration — exactly what you want on a runner.

The JSON contract

Headless runs support three output formats: text (default), json (a single result envelope), and stream-json (newline-delimited events). For pipelines, json is the workhorse.

The envelope is a single object carrying the result and its operational metadata. You get result (the text), subtype (success or error subtypes like error_max_turns), is_error, session_id, num_turns, and duration_ms. You also get total_cost_usd, usage token counts, and a per-model modelUsage breakdown — the docs explicitly pitch total_cost_usd for tracking spend per invocation.

For typed output, add --json-schema with a JSON Schema. The structured payload lands in a structured_output field alongside the metadata. Here is the release-notes primitive the third recipe builds on:

git log --oneline v2.3.0..HEAD | claude --bare -p \
    "Categorize these commits into features, fixes, and breaking \
changes for release notes. Use the exact commit subjects." \
    --output-format json \
    --json-schema '{"type":"object","properties":{"features":
{"type":"array","items":{"type":"string"}},"fixes":{"type":"array",
"items":{"type":"string"}},"breaking":{"type":"array","items":
{"type":"string"}}},"required":["features","fixes","breaking"]}' \
    | jq '.structured_output'

The jq filter prints just the typed object: three arrays, features, fixes, and breaking, each holding commit subjects. Drop the filter and you get the whole envelope, with that object under structured_output next to session_id, total_cost_usd, usage, and the rest of the metadata above. If the value you pass isn’t a valid JSON Schema, claude exits with an error before the run starts. Before v2.1.205 it silently ignored an invalid schema and returned plain text, which is why we pin a floor version for any job that depends on the schema.

The payoff: the next pipeline step consumes structured_output with jq instead of regex-parsing prose. Without a schema, extract the plain text with jq -r '.result'.

Pro tip

Pipe data in via stdin instead of granting read permissions. When the input arrives on stdin, the run needs no Bash or file tools to see it. The official docs use this exact pattern for a least-privilege PR security review.

Here is that documented pattern, verbatim:

gh pr diff "$1" | claude -p \
  --append-system-prompt \
  "You are a security engineer. Review for vulnerabilities." \
  --output-format json

The diff is input, not something Claude fetches, so the run gets zero Bash and zero network access. Add --bare and --max-turns 6, then post jq -r '.result' back with gh pr comment.

Use stream-json when you need progress and logs rather than a single envelope. Its system/init event lists loaded plugins and a plugin_errors array — the docs suggest failing CI when that array is non-empty. It also emits system/api_retry events you can surface in job logs.

Permissions when nobody is watching

Interactive sessions have a human to approve tool calls. Headless runs need a policy instead. Four modes matter:

ModeBehavior in -p runs
dontAskAuto-denies anything that would prompt. Only permissions.allow rules and read-only commands execute.
acceptEditsAuto-approves file edits and common filesystem commands in the working directory. Other commands still need allow rules or the run aborts.
bypassPermissionsEverything runs. Containers and VMs only.
autoClassifier-checked. In -p, once repeated blocks hit the threshold, blocked actions are skipped and Claude keeps working, so the outcome is harder to predict — a poor fit for cron.

Default to --permission-mode dontAsk and widen only with a reason. Unattended, an over-permissioned agent can push, delete, or call out to the network with nobody to stop it.

For any job with no human on the other end, the docs now recommend adding --permission-prompts none (v2.1.259 or later). Anything that would prompt is denied unless a PermissionRequest hook allows it, Claude is told nobody can approve it and not to retry, and tools that need a person, such as AskUserQuestion, are removed. It matters most when a permission host is attached, such as an Agent SDK canUseTool callback or --permission-prompt-tool, because without it the run waits for that host to answer.

Grant capabilities with --allowedTools using permission rule syntax. This documented example scopes a run to exactly the git commands a commit job needs:

claude -p "Look at my staged changes and create an appropriate commit" \
  --allowedTools \
  "Bash(git diff *),Bash(git log *),Bash(git status *),Bash(git commit *)"
Warning

The space before * matters in prefix rules. Bash(git diff *) matches git diff --stat but not git diff-index. Bash(git diff*) — no space — matches both. One character silently widens the rule to commands you never reviewed.

Deny rules block in every mode, including bypassPermissions. Writes to protected paths — .git, .claude, .mcp.json, shell rc files, pre-commit hook configs, and more — are never auto-approved outside bypassPermissions, and an allow rule in settings does not change that; in dontAsk they are denied outright.

For enforcement beyond permission rules, hooks are the hard layer underneath. Remember that --bare skips hook auto-discovery, so pass CI hooks explicitly via --settings.

Danger

bypassPermissions (alias: --dangerously-skip-permissions) executes everything, including protected-path writes. The docs are blunt: only use it in isolated environments like containers or VMs without internet access — harden the sandbox first. On Linux and macOS it refuses to start as root outside a recognized sandbox, and an rm of a critical path such as rm -rf / still asks for approval as a circuit breaker. But in headless mode, a prompt with nobody watching is an abort, not a save.

Guardrails for unattended runs

Claude Code headless mode turns the agent into a service, so operate it like one. Treat it like any flaky network dependency: bound it, watch it, and assume it will occasionally misbehave.

Bound turns and wall clock. --max-turns caps the agent loop; a pipeline timeout caps real time. Always pair them:

timeout 600 claude --bare -p "..." --max-turns 10 --output-format json

Turn limits stop runaway loops. Timeouts stop a single turn that hangs on a slow command.

Branch on exit codes, diagnose from the envelope. Treat any non-zero exit as failure rather than assuming specific codes. Then read is_error and subtype from the JSON to learn why — error_max_turns means your budget was too small, not that the task is impossible.

Prevent overlap. Cron does not wait for the previous run. Use flock in crontab, or a concurrency group in GitHub Actions:

0 3 * * 1-5 flock -n /tmp/claude-deps.lock -c \
  'cd /srv/repo && timeout 1200 claude --bare -p "$(cat prompt.md)" \
   --output-format json >> /var/log/claude-deps.jsonl'

Appending each envelope to a JSONL file also gives you a per-run audit log for free. For fleet-level metrics, export OTEL data with CLAUDE_CODE_ENABLE_TELEMETRY=1.

Write closed-ended prompts. Nobody answers questions in a cron job. A prompt that invites follow-up is a bug:

Before
Bump the outdated packages and let me know if you also want the tests updated.
After
Bump minor and patch versions only. Run npm test. If green, open a PR and stop. If red, revert and report the failure. Ask nothing.

Every headless prompt needs explicit inputs, a done-condition, and output instructions. A headless prompt is a spec by necessity — there is no conversation to recover in.

Two operational limits worth knowing. Claude Code caps piped stdin at 10MB — exceeding it exits non-zero, so write big logs to a file and reference the path. Claude Code also terminates background Bash tasks about 5 seconds after a -p run’s final result.

Recipe 1: triage failing CI runs

The read-only recipe. When CI fails, a triage job reads the failed log, inspects the source, and posts a structured diagnosis to the PR. It can read everything and change nothing.

Define the output contract first:

{
  "type": "object",
  "properties": {
    "culprit_area": { "type": "string" },
    "likely_cause": { "type": "string" },
    "is_flaky": { "type": "boolean" },
    "suggested_fix": { "type": "string" },
    "confidence": { "enum": ["low", "medium", "high"] }
  },
  "required": ["culprit_area", "likely_cause", "is_flaky",
               "suggested_fix", "confidence"]
}

Then the workflow:

name: CI failure triage
on:
  workflow_run:
    workflows: ["CI"]
    types: [completed]

permissions:
  contents: read
  actions: read
  pull-requests: write

jobs:
  triage:
    if: github.event.workflow_run.conclusion == 'failure'
    runs-on: ubuntu-latest
    concurrency:
      group: triage-${{ github.event.workflow_run.head_branch }}
      cancel-in-progress: true
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.workflow_run.head_sha }}

      - run: npm install -g @anthropic-ai/claude-code

      - name: Fetch the failing log
        env:
          GH_TOKEN: ${{ github.token }}
        run: |
          gh run view ${{ github.event.workflow_run.id }} \
            --log-failed > failed.log

      - name: Triage
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: |
          PROMPT='Read failed.log at the repo root. Identify the
          failing step and the most likely cause. Decide whether the
          failure looks flaky. Inspect source files if needed, but do
          not modify anything. Fill every schema field, then stop.'
          timeout 300 claude --bare -p "$PROMPT" \
            --permission-mode dontAsk \
            --allowedTools "Read,Grep,Glob" \
            --max-turns 6 \
            --output-format json \
            --json-schema "$(cat .github/triage-schema.json)" \
            > triage.json
          jq -e '.is_error == false' triage.json

      - name: Comment on the PR
        if: github.event.workflow_run.pull_requests[0]
        env:
          GH_TOKEN: ${{ github.token }}
        run: |
          jq -r '.structured_output |
            "**CI triage**\n\n" +
            "- Area: \(.culprit_area)\n" +
            "- Likely cause: \(.likely_cause)\n" +
            "- Flaky: \(.is_flaky)\n" +
            "- Suggested fix: \(.suggested_fix)\n" +
            "- Confidence: \(.confidence)"' triage.json > comment.md
          gh pr comment \
            ${{ github.event.workflow_run.pull_requests[0].number }} \
            --body-file comment.md

Note the shape of the guardrails: dontAsk plus a read-only tool list, six turns, a five-minute timeout, and a jq -e gate that fails the step if the envelope reports an error. The log goes through a file, not stdin, because failed CI logs routinely blow past the 10MB stdin cap.

Want an escalation path? Chain a second stage off the same session, gated on confidence == "high":

session_id=$(jq -r '.session_id' triage.json)
claude -p "Attempt the suggested fix on a new branch. Run the
failing test. Push the branch only if it passes." \
  --resume "$session_id" \
  --permission-mode acceptEdits \
  --allowedTools "Bash(npm test),Bash(git checkout -b *),Bash(git push *)"

Since v2.1.223, --resume <id> finds the session in any project on the machine, so session lookup no longer ties the two stages to one directory. The workspace still matters for everything else: the second stage needs the same checkout and the same ~/.claude session store, so run both stages as steps in one job rather than as separate jobs on fresh runners. When confidence is low, hand off to a human-paced session instead: triage is the read-only half, and babysitting the PR to green is the interactive follow-up.

Recipe 2: nightly dependency PRs

The write-capable recipe, and the one that beats Dependabot on one axis: the agent runs and interprets your test suite before the PR exists.

  1. 01

    Schedule the run with overlap protection

    Use a schedule trigger and a concurrency group. On a plain server, use the crontab-plus-flock pattern from the guardrails section instead.

    name: Nightly dependency updates
    on:
      schedule:
        - cron: "0 3 * * 1-5"
    concurrency:
      group: nightly-deps
    
    jobs:
      bump:
        runs-on: ubuntu-latest
        permissions:
          contents: write
          pull-requests: write
        steps:
          - uses: actions/checkout@v4
          - run: npm install -g @anthropic-ai/claude-code
  2. 02

    Scope the permissions to the exact commands

    Build the allow list from prefix rules — nothing broader than the job needs. Mind the space before each *.

          - name: Bump one cohort
            env:
              ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
              GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
            run: |
              TOOLS="Bash(npm outdated *),Bash(npm install *)"
              TOOLS="$TOOLS,Bash(npm test),Bash(git checkout -b *)"
              TOOLS="$TOOLS,Bash(git add *),Bash(git commit *)"
              TOOLS="$TOOLS,Bash(git push *),Bash(gh pr create *)"
              timeout 1200 claude --bare -p \
                "$(cat .github/prompts/nightly-deps.md)" \
                --permission-mode acceptEdits \
                --permission-prompts none \
                --allowedTools "$TOOLS" \
                --max-turns 30 \
                --output-format json > run.json
              jq -e '.is_error == false' run.json

    There is no gh pr merge rule and no git push to main in the prompt. The agent cannot merge its own work through Claude Code’s permissions. --permission-prompts none makes the unattended intent explicit: anything outside the allow list is denied, and Claude is told not to retry it.

  3. 03

    Write a closed-ended prompt with a failure branch

    Keep the prompt in a versioned file so changes go through review.

    Update npm dependencies. Follow these steps exactly:
    
    1. Run npm outdated. Select only minor and patch updates.
    2. Create a branch named deps/nightly-<today's date>.
    3. Install the updated versions.
    4. Run npm test.
    5. If tests pass: commit, push the branch, and open a PR with
       gh pr create. List each bump and the test result in the body.
    6. If tests fail: identify the offending package, revert it, and
       rerun the tests. If they still fail, stop and report the
       failure in your final message. Do not open a PR.
    
    Never merge the PR. Never push to main. Stop when the PR is open
    or when you have reported a failure.
  4. 04

    Let the test suite gate the PR

    Step 4 of the prompt is the heart of the recipe: the suite runs inside the agent loop, and a red suite means no PR. This is a verification loop running unattended — the agent’s claim of success is checked by your tests, not taken on faith.

  5. 05

    Keep the merge human

    The pipeline opens the PR; a person merges it. That person needs a real review discipline for agent-generated diffs — see reviewing AI diffs without rubber-stamping. Check total_cost_usd in run.json too; a nightly job’s spend should be boring and flat.

Recipe 3: release notes from a diff

The zero-write recipe. On a tag push, pipe the commit log in via stdin, get typed categories out, and render them into the GitHub release. Claude needs no repo permissions at all — the data arrives as input.

name: Release notes
on:
  push:
    tags: ["v*"]

jobs:
  notes:
    runs-on: ubuntu-latest
    permissions:
      contents: write
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - run: npm install -g @anthropic-ai/claude-code

      - name: Generate notes
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: |
          prev=$(git describe --tags --abbrev=0 "$GITHUB_REF_NAME^")
          PROMPT='Categorize these commits into features, fixes, and
          breaking changes for release notes. Use the exact commit
          subjects.'
          git log --oneline "$prev..$GITHUB_REF_NAME" | \
            timeout 300 claude --bare -p "$PROMPT" \
            --permission-mode dontAsk \
            --max-turns 3 \
            --output-format json \
            --json-schema "$(cat .github/release-schema.json)" \
            > notes.json
          jq -e '.is_error == false' notes.json

      - name: Publish release
        env:
          GH_TOKEN: ${{ github.token }}
        run: |
          jq -r '.structured_output |
            "## Features\n" +
            (.features | map("- \(.)") | join("\n")) +
            "\n\n## Fixes\n" +
            (.fixes | map("- \(.)") | join("\n")) +
            "\n\n## Breaking changes\n" +
            (.breaking | map("- \(.)") | join("\n"))' \
            notes.json > notes.md
          gh release create "$GITHUB_REF_NAME" --notes-file notes.md

The schema file is the three-array structure from the JSON contract section (features, fixes, breaking, all required). Three turns is plenty for a categorization task with all input on stdin — if this job ever hits error_max_turns, something upstream changed.

When not to run your own scheduler

Anthropic now ships a managed alternative: routines. A routine is a saved Claude Code configuration — a prompt, repositories, and connectors — that runs on a schedule, from an API call, or on GitHub events, on Anthropic-managed cloud infrastructure (or your organization’s self-hosted environment). As of this writing, routines are in research preview on Pro, Max, Team, and Enterprise plans, created at claude.ai/code/routines or with /schedule from the CLI. Check the routines docs for current limits.

The DIY route in this guide wins when you need your own runners, secrets, and OIDC. It also wins for provider choice (Bedrock or Vertex), build-artifact access, and audit logs inside your own stack. Ownership is the other difference: a routine belongs to one person’s claude.ai account, and its commits and pull requests carry that person’s GitHub identity. A CI job runs as whatever service identity you give it.

If none of that applies, a routine may be less infrastructure for the same outcome. For recurring checks inside a live session, /loop is lighter still. For a devops team already operating CI, Claude Code headless mode slots into what you run today.

The recipes are the easy part. What keeps them running for months is the same discipline as any other production service: a named owner per job, spend that someone watches, permission lists that change only through review, and a clear line where the agent stops and a person merges. Putting that operating model around Claude jobs is the core of how we approach taking Claude systems to production.


Next steps

External references: the official headless mode docs and permission modes reference cover every flag used here.

Related service

This is part of our Build work — make it survive production.