Claude Code OpenTelemetry metrics, decoded

Turn Claude Code OpenTelemetry metrics and local JSONL transcripts into cost-per-task, interrupt-rate, and permission-friction data you act on.

on this page

“It feels slow” is not a bug report. Neither is “it feels expensive this month” or “I keep having to step in.” Claude Code OpenTelemetry metrics turn those feelings into numbers — and the JSONL transcripts already sitting in ~/.claude/projects/ hold months of history you can mine without deploying anything.

We ask the same four questions of every Claude Code workflow we look at:

  1. Cost per task, not cost per month — sliced by model, subagent, skill, and MCP server.
  2. How often you interrupt — every Esc and “No” is a correction, and your correction rate is a quality signal.
  3. Where permission prompts break flow — which tools stall you, and for how long.
  4. Which instructions get ignored — repeated corrections point at CLAUDE.md rules that are not landing.

This is not a dashboard tutorial — the dashboard is the means. The end is a loop: measure, rank the offenders, fix the worst one, re-measure. Measurement without a fix loop is dashboard theater.

What you’ll learn

  • Mine local JSONL transcripts with jq for cost, tool mix, and interrupt rate — retroactively, with zero infrastructure
  • Enable OpenTelemetry export with the right exporters, endpoints, and privacy gates
  • Map documented metrics, events, and beta trace spans to specific workflow defects
  • Quantify permission friction with tool_decision sources and blocked_on_user span time
  • Close the loop: ship one config fix, then prove it worked with a re-measure

Prerequisites

  • Claude Code CLI with weeks of real usage on a persistent machine. Transcripts do not survive ephemeral containers; those environments need the OpenTelemetry path.
  • jq installed.
  • For the export path: Docker, plus working knowledge of OTLP, Prometheus, and Grafana. This guide does not explain OpenTelemetry concepts.

Start with the transcripts you already have

Every session appends to ~/.claude/projects/<encoded-project-path>/<session-id>.jsonl. Each line is a typed record — user, assistant, system, plus bookkeeping types like queue-operation. Filter by type; never assume every line is a message.

Assistant lines carry message.model and a message.usage object with input_tokens, output_tokens, cache_creation_input_tokens, and cache_read_input_tokens. Tool calls appear as tool_use content blocks inside assistant messages. Results come back as tool_result blocks in user-typed lines, matched by tool_use_id.

That schema answers the cost and tool-mix questions for every session you have ever run — history OpenTelemetry never saw. Paste this prompt into any project:

Analyze my Claude Code transcripts in ~/.claude/projects/.
For each project directory:
1. Sum token usage across assistant turns (input, output,
   cache_read, cache_creation)
2. Count tool_use blocks by tool name
3. Count lines containing "[Request interrupted by user]" and
   tool_result blocks containing "User rejected tool use"
Give me a table per project and flag the session with the worst
interrupt count. Use jq, don't read the files into context.

Claude Code reaches for jq instead of reading megabytes of JSONL. The core commands look like this:

cd ~/.claude/projects && jq -r 'select(.type=="assistant")
  | .message.usage | [.input_tokens, .output_tokens,
  .cache_read_input_tokens, .cache_creation_input_tokens]
  | @tsv' -- */*.jsonl |
  awk -F'\t' '{i+=$1; o+=$2; cr+=$3; cc+=$4; n++} END
  {printf "turns=%d in=%d out=%d cr=%d cc=%d\n", n, i, o, cr, cc}'

jq -r 'select(.type=="assistant") | .message.content[]?
  | select(.type=="tool_use") | .name' -- */*.jsonl |
  sort | uniq -c | sort -rn

The first prints one line of totals: assistant turns, fresh input, output, cache reads, and cache writes. The second prints a count per tool name, most-used first.

Read the ratio before the totals. On a machine with weeks of history, we expect cache reads to dwarf fresh input — that cache-read column is what keeps long sessions affordable. Watching it collapse right after a compaction is measurable evidence of the context thrash that surviving auto-compact teaches you to avoid.

Pro tip

Ask Claude Code to jq its own transcripts instead of reading the JSONL into context. Transcript analysis is exactly the kind of task that would otherwise blow the context window — and the one-liners it writes become your permanent analysis scripts.

If you only want the cost answer, skip the scripting. npx ccusage@latest daily reads the same JSONL and prints daily token and estimated-cost tables per model, no telemetry setup required. ccusage session is the closest off-the-shelf approximation of cost per task, and --json makes it scriptable for threshold alerts.

Mine for behavior, not just tokens

Token totals tell you what Claude Code spent. Transcripts also record how often you had to stop it, and that number is the better quality signal.

Interruptions appear as the literal text [Request interrupted by user]. Rejected tools land as tool_result blocks containing “User rejected tool use”. Count both per file:

cd ~/.claude/projects
grep -rc --include='*.jsonl' '\[Request interrupted by user\]' .
grep -rc --include='*.jsonl' 'User rejected tool use' .

One caveat: the marker records that a turn was cut short, not why, so treat the count as an upper bound. Even so, sessions with the worst interrupt counts are where you paid the correction tax — this grep is that tax, measured.

The fourth question — which instructions get ignored — has no telemetry signal at all. It lives only in your own user messages:

Search my user messages in ~/.claude/projects/**/*.jsonl for
correction phrases: "no, use", "I said", "again", "stop", "don't".
Cluster them by topic and show the three corrections I repeat most
across sessions. Use jq and grep, don't read files into context.

If you correct the same convention three times across sessions, that is a failing contract clause. The rule is either missing from CLAUDE.md or written in a way the model skips — CLAUDE.md as a contract, not a wiki covers the rewrite.

Turn on Claude Code OpenTelemetry metrics

Transcripts answer questions retroactively. Claude Code OpenTelemetry metrics answer them continuously, with structure the JSONL lacks: per-request cost, permission decisions with their sources, and compaction events. A prompt.id ties every event back to the user prompt that caused it.

Export is opt-in and speaks standard OTLP to your own collector.

Note

OpenTelemetry export is separate from Anthropic’s product telemetry; data goes only to the endpoint you configure. Content is redacted by default — prompts, responses, and tool parameters stay out unless you set OTEL_LOG_USER_PROMPTS, OTEL_LOG_ASSISTANT_RESPONSES, or OTEL_LOG_TOOL_DETAILS.

The export does include user.email and account IDs by default when authenticated. Fine for a personal collector; a policy decision for a shared org one.

  1. 01

    Enable telemetry in settings.json

    Set the env vars under "env" in ~/.claude/settings.json so telemetry is on for every session, with no shell-profile hacks:

    {
      "env": {
        "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
        "OTEL_METRICS_EXPORTER": "console",
        "OTEL_METRIC_EXPORT_INTERVAL": "10000"
      }
    }

    The default export interval is 60000 ms. Drop it to 10000 while debugging so you are not staring at a silent terminal for a minute.

  2. 02

    Smoke-test with the console exporter

    Start a session and watch the terminal. Metric batches print every 10 seconds: expect claude_code.token.usage with attributes type (input, output, cacheRead, cacheCreation), model, and session.id, plus claude_code.cost.usage in USD.

    In our experience, most “no data in Grafana” problems are exporter or backend misconfiguration, not Claude Code. The console exporter proves data flows before any backend enters the picture.

  3. 03

    Point at a collector

    Swap the exporters to OTLP:

    {
      "env": {
        "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
        "OTEL_METRICS_EXPORTER": "otlp",
        "OTEL_LOGS_EXPORTER": "otlp",
        "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
        "OTEL_EXPORTER_OTLP_ENDPOINT": "http://localhost:4317"
      }
    }

    For the backend, start from Anthropic’s claude-code-monitoring-guide repo, which ships Docker Compose, OpenTelemetry collector, and Prometheus configurations. We start there rather than writing a compose file from scratch.

  4. 04

    Verify data lands

    Open Grafana, or whatever sits on top of your backend, and confirm the cost and token panels populate during a live session. Then set OTEL_METRIC_EXPORT_INTERVAL back to the default.

Warning

Claude Code exports metrics with delta temporality by default, and a backend that expects cumulative temporality may not ingest delta datapoints. If panels stay empty after the console exporter worked, set OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative, and remember Prometheus-style backends translate the dots in claude_code.token.usage to underscores. In our experience, temporality mismatch is the most common cause of an empty dashboard.

Four views that answer the four questions

The metric catalog is larger than you need. Four rollups map straight to the four questions.

Cost per task, not per month

claude_code.cost.usage and claude_code.token.usage carry the attributes that make slicing free: model, query_source (main, subagent, auxiliary), agent.name, skill.name, mcp_server.name, and mcp_tool.name. “What does my code-review subagent cost per week” is a label query, not custom instrumentation. One catch: user-defined agent names and user-configured MCP server names are exported as custom unless you also set OTEL_LOG_TOOL_DETAILS=1, so decide whether that detail belongs in your collector before you build panels on it.

For true per-prompt cost, use the claude_code.api_request event instead. Each carries an estimated cost_usd, duration_ms, and token counts, and every event descending from one user prompt shares a prompt.id. Group by prompt.id and you have cost per task.

Two slices earn a panel immediately. mcp_server.name shows which server’s tool results burn the most tokens — the evidence you want before cutting servers to protect your context budget. And query_source=subagent with agent.name verifies a subagent actually firewalls context rather than just relocating the spend.

Interrupt and rejection rate

The claude_code.tool_decision event fires on every permission decision with tool_name, decision, and a source that does the diagnostic work:

  • config — auto-decided by your settings rules. Free.
  • hook — decided by a hook. Free, and visible: hooks as guardrails show up here as decisions that never cost you a prompt.
  • user_permanent — you chose “don’t ask again.”
  • user_temporary — you sat through a prompt and said yes once.
  • user_reject and user_abort — you said no, or dismissed the prompt.

user_reject plus user_abort is your live interruption rate, the forward-looking twin of the transcript grep. For edits specifically, the claude_code.code_edit_tool.decision metric tracks accept versus reject on Edit and Write.

Permission friction

Every user_temporary decision is a prompt you sat through that an allowlist rule or hook could have absorbed. The ratio of user_temporary to config is your friction score.

Beta traces make friction literal. With CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 and OTEL_TRACES_EXPORTER=otlp, the claude_code.tool.blocked_on_user span records how long each prompt held the session, with decision and source attached. Sum it per day for “minutes lost to permission prompts”; group by the parent span’s tool_name to know what to allowlist first.

Traces are beta, so span names and attributes may change. The sibling llm_request spans also record ttft_ms per model — the data that separates “the model is slow today” from “my context is bloated.”

Context health

The claude_code.compaction event logs trigger (auto or manual), pre_tokens, and post_tokens. Frequent auto-compactions plus a collapsing cacheRead share of token.usage means context thrash: you are paying to rebuild cache that compaction destroyed.

That pair of signals is the measurable side of context budgeting. Watch it weekly; it degrades quietly as CLAUDE.md and MCP servers accrete. When a session’s window is already blown, surviving auto-compact covers triage.

Fix the top offender, then re-measure

Numbers without a config change are dashboard theater. Rank your offenders, then map each signal to its fix:

SignalDefectFix
High user_temporary count on one toolMissing allowlist ruleAdd permissions.allow rule or a hook
Long blocked_on_user totalsWrong permission postureAllowlist; check permission_mode_changed flapping
High interrupt rate on one task typeUnder-specified promptsPlan first; tighten the spec
Same correction three-plus timesBroken CLAUDE.md clauseRewrite as contract rule or hook
Frequent trigger=auto compactionsContext budget blownTrim CLAUDE.md; scope the session

Two of those fixes lean on hooks; hooks as guardrails covers how we write them.

Make it concrete. Say a week of data shows 41 user_temporary decisions, 29 of them on Bash running your test and lint commands. Convert exactly those into allowlist rules in .claude/settings.json:

Before
"allow": []
After
"allow": ["Bash(npx vitest run *)", "Bash(npm run lint *)"]

Re-measure the following week. user_temporary should drop toward zero for those commands while config rises — same work, no waiting. If your top offender is interrupt rate instead, the fix is upstream in plan mode or a tighter spec, not in permissions.

Then hold the cadence: a 15-minute weekly review. One query per question, one worst offender, one config change, and a comparison against last week’s numbers. One fix per week compounds; five at once tells you nothing about which one worked.

This loop is also how we start any production engagement. Before we recommend a change to a team’s Claude Code setup, we want a baseline for cost per task, permission friction, and interrupt rate, because without one nobody can tell whether the change helped. That baseline is the first thing an Assess engagement produces, and the re-measure is how we show a fix actually worked.

Next steps

  • The fix-and-re-measure cadence is a verification loop pointed at your own workflow instead of Claude’s code.
  • Cost-per-task data tells you which rung of the escalation ladder each task class belongs on.
  • Instrument headless Claude Code in CI the same way — set OTEL_METRICS_INCLUDE_ENTRYPOINT=true and the app.entrypoint attribute separates CI and SDK sessions from interactive ones.
  • For the full metric and event reference, keep the official monitoring docs open; the monitoring-guide repo ships the collector and backend configurations.

Related service

This is part of our Assess work — decide what to build and why.