Document extraction that knows when to ask a person

An event-driven pipeline that extracts structured data from documents with Claude on Amazon Bedrock, checks it against evidence, and sends only the doubtful cases to human review.

on this page

the problem

Invoices, claims and contracts arrive by the thousand, and today someone keys the fields out of them by hand: supplier, invoice number, date, amount, parties.

The business wants most of those documents processed without a person touching them. But a wrong amount or a wrong party is expensive. So the real requirement isn’t “extract fields”. It is: extract fields, know which ones you can’t vouch for, and send only those to a person.

Around that sit the non-functionals we design for from the start:

  • Throughput and bursts. Documents often arrive in batches, at month end for example, not evenly through the day.
  • PII. Names, addresses, bank details and ID numbers are often the very fields we’re extracting.
  • An audit trail per document. When a figure is questioned, someone will ask what the system read, what it decided and who changed it.
  • Cost per document. OCR, model tokens and reviewer minutes all add up.

the design

Accent marks where evidence decides and where a person decides.
  1. A document is uploaded to the inbox bucket in Amazon S3, encrypted with KMS.
  2. S3 sends an Object Created event to Amazon EventBridge (EventBridge notifications must be enabled on the bucket), and a rule starts a Step Functions workflow.
  3. The workflow’s first step writes an idempotency key to DynamoDB with a conditional write. If the key already exists, this is a duplicate and the execution stops.
  4. When the document needs it, Textract or Bedrock Data Automation reads it and returns text, layout and per-block confidence. Born-digital, text-only PDFs skip this step.
  5. Claude on Amazon Bedrock extracts the fields with structured output, constrained to a JSON schema.
  6. Code checks the result: validators, business rules and grounding in the source. On failure, the errors go back to Claude for one retry, and the result is checked again.
  7. A Choice state routes documents that pass every check to be written to the system of record.
  8. Everything else goes to the review queue. The workflow sends a task token through Amazon SQS and waits on it, a reviewer corrects or confirms the fields in the review UI, and the UI resumes the workflow.
  9. Every outcome, accepted or reviewed, writes an audit record and the eval data that comes from it.

The review wait uses .waitForTaskToken, which can hold an execution open for up to the one-year Step Functions quota. That is plenty for a queue that empties over a weekend, and it means we use a Standard workflow, not an Express one.

The extraction call uses the Converse outputConfig. Here is a trimmed invoice schema. Every extracted field is nullable, and additionalProperties is false, as the schema subset requires:

{
  "type": "object",
  "properties": {
    "supplier_name": { "type": ["string", "null"] },
    "invoice_number": { "type": ["string", "null"] },
    "total_amount": { "type": ["string", "null"], "description": "As printed on the document" },
    "not_found": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": { "field": { "type": "string" }, "reason": { "type": "string" } },
        "required": ["field", "reason"],
        "additionalProperties": false
      }
    }
  },
  "required": ["supplier_name", "invoice_number", "total_amount", "not_found"],
  "additionalProperties": false
}

The schema goes into outputConfig as a JSON string:

{
  "textFormat": {
    "type": "json_schema",
    "structure": {
      "jsonSchema": {
        "schema": "<the schema above, serialised as a string>",
        "name": "invoice_extraction"
      }
    }
  }
}

The same schema works as a strict tool’s input schema, with toolChoice set to that tool.

The check step first checks the stop reason, then validates in order: schema, then business rules, then grounding. On failure it sends one follow-up turn with the error list, and then gives up to review:

import json
import boto3
from jsonschema import Draft202012Validator

bedrock = boto3.client("bedrock-runtime")

def extract(model_id, messages, output_config, schema, checks):
    """checks: functions taking the result and returning a list of error strings,
    business rules first, then grounding against the document's text."""
    validator = Draft202012Validator(schema)
    for attempt in range(2):  # the first pass plus one retry
        resp = bedrock.converse(modelId=model_id, messages=messages, outputConfig=output_config)
        if resp["stopReason"] != "end_turn":  # refusal or token limit: output may not match
            return {"route": "review", "result": None, "errors": [f"stop reason {resp['stopReason']}"]}
        reply = resp["output"]["message"]
        result = json.loads(reply["content"][0]["text"])
        errors = [e.message for e in validator.iter_errors(result)]
        if not errors:
            errors = [err for check in checks for err in check(result)]
        if not errors:
            return {"route": "accept", "result": result, "attempts": attempt + 1}
        messages = messages + [reply, {"role": "user", "content": [{"text": (
            "These values failed our checks. Correct them from the document, "
            "or set them to null and list them in not_found:\n" + "\n".join(errors))}]}]
    return {"route": "review", "result": result, "errors": errors}

If you use strict tool use instead, the follow-up is a toolResult with status: "error" carrying the same list.

In the workflow, a Choice state reads the route, and anything not accepted goes to a review task that sends the task token to the queue and waits (trimmed):

"Route": {
  "Type": "Choice",
  "Choices": [
    { "Variable": "$.check.route", "StringEquals": "accept", "Next": "WriteRecord" }
  ],
  "Default": "Review"
},
"Review": {
  "Type": "Task",
  "Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken",
  "Parameters": {
    "QueueUrl": "https://sqs.<region>.amazonaws.com/<account-id>/extraction-review",
    "MessageBody": {
      "documentId.$": "$.documentId",
      "reason.$": "$.check.errors",
      "taskToken.$": "$$.Task.Token"
    }
  },
  "ResultPath": "$.review",
  "Next": "WriteRecord"
}

A small consumer turns each message into a row in the task table. When the reviewer submits, the UI calls SendTaskSuccess with that token and the corrected fields, and the workflow carries on with them.

components

componentAWS serviceresponsibility
InboxAmazon S3 (KMS-encrypted)Receives documents, with lifecycle rules on raw files.
EventAmazon EventBridgeMatches Object Created events for the inbox and starts the workflow.
DedupeAmazon DynamoDBHolds idempotency keys, written with a conditional write so each document is processed once.
WorkflowAWS Step FunctionsOrchestrates read, extract, check and route, and owns throttling retries.
ReadAmazon Textract or Bedrock Data AutomationOCR and layout for scanned or layout-heavy documents, with per-block confidence.
ExtractClaude on Amazon BedrockFills the schema from the document’s text and layout. Has no tools with side effects.
Checkyour code (Lambda or a workflow step)Validators, business rules and grounding. Produces the routing decision and its reason.
RouteStep Functions Choice stateSends each document to accept or to review, based on the check results.
Acceptyour system of recordReceives the accepted fields.
Review queuea queue we own: Amazon SQS plus a DynamoDB task tableThe workflow sends each waiting document’s task token and review reason to SQS; a small consumer writes it to the task table the review UI reads.
Revieweryour review UIShows the document and the flagged fields, and calls SendTaskSuccess with the corrected values.
Audit and eval dataAmazon S3 and DynamoDBOne record per document, plus reviewer corrections as labelled examples.

Amazon A2I is in maintenance mode and closed to new customers since July 2026, so we don’t build new pipelines on it. The queue we own is a small amount of code, and it keeps the review data where the evals can use it.

decisions

How Claude sees the document

  • Converse with a PDF document block
  • Chosen Textract or Bedrock Data Automation first, then Claude on the text and layout
  • Page images
  • InvokeModel with a document block

On Converse, a PDF document block is read as text only unless citations are on, and citations can’t be combined with structured outputs. With the schema we want, a scanned page or a table’s layout would be lost.

Running OCR and layout extraction first gives Claude the text and its layout, plus a per-block confidence from Textract that the grounding check uses later. For born-digital, text-only documents, skip this step and send the PDF directly.

Page images and InvokeModel with a document block both give Claude the visual page, and are worth testing for layout-heavy documents. The trade-off of our choice is one more service and one more cost line.

How the output shape is enforced

  • Ask for JSON in the prompt
  • Strict tool use
  • Chosen Structured output with a JSON schema

Constrained decoding guarantees the result parses against the schema on a normal stop, so no more repair prompts for broken JSON. A refusal or a stop at the token limit can still return output that doesn’t match, so the code checks the stop reason first.

Give each field type: ["string", "null"] so the model can say “not found” instead of inventing a value, and add a not_found list for the reason. The schema subset can’t express ranges or lengths, so those checks live in code. Strict tool use is equivalent if the extraction is one step of a larger tool-using flow.

The schema guarantees the shape of the answer, not that it is true.

Where the routing signal comes from

  • The model's own confidence score
  • Chosen Evidence: validators, business rules, grounding in the source, agreement between passes

We route on four kinds of evidence:

  • deterministic validators, such as formats and check digits;
  • business rules: totals add up, dates are in range, the supplier exists;
  • a grounding check that each value appears in the source text, or carries a Textract or Bedrock Data Automation confidence;
  • agreement between two passes, for high-consequence fields only.

A required field that comes back null is a signal too. A confidence number the model reports about itself is at most a tie-breaker, never a gate.

How many retries before review

  • Retry until valid
  • Chosen One retry with the errors, then review
  • No retry

One follow-up turn that feeds back the specific validation errors fixes most slips. Beyond that, retries cost money and hide real ambiguity, which is exactly what a person should see.

Throttling and timeouts are a different kind of failure. They are handled by the Step Functions Retry with backoff, not by this budget.

Who sets the review threshold

  • Engineering picks a number
  • Chosen The business owner sets it per field and consequence, starting with everything in review

How much error is acceptable is a business decision. Amount fields and party names deserve stricter rules than a reference field.

Start with every document reviewed. Then let auto-accept earn its way in, field by field, as review outcomes show the checks are reliable. Name the owner, and write their thresholds down next to the rules they govern.

How duplicates are stopped

  • Trust S3 to deliver each event once
  • Chosen Idempotency key with a conditional write, plus business-level duplicate checks

S3 event notifications are at-least-once, and duplicates carry the same key, version ID and sequencer.

Key on bucket, key and version ID, or on a SHA-256 of the content, and record it with a DynamoDB conditional write. Where you start executions yourself, use the same key as the Step Functions execution name, hashed to fit the naming rules, so Step Functions won’t start a second execution for the same document.

Separately, catch the same invoice arriving twice under different file names, for example the same supplier and invoice number. That check belongs with the business rules, and a hit goes to review.

safety and governance

The document is untrusted input. Anyone who can send an invoice can put text in it that reads like instructions. So:

  • it goes in only as document or text content with a neutral name, never as instructions, and never in the system prompt;
  • the extraction step has no tools with side effects, so the worst a malicious document can do is plant a wrong value;
  • validators and grounding are there to catch planted values.

For PII:

  • KMS on the bucket, and PrivateLink to Bedrock so traffic stays off the public internet;
  • tag PII fields with Comprehend DetectPiiEntities, or Guardrails in Detect mode, but don’t mask before extraction when the PII is the field;
  • mask PII in logs and in the reviewer UI unless the reviewer’s role needs it;
  • lifecycle rules on raw documents, matched to your retention policy;
  • restrict model-invocation logging, which, when enabled, can store the full text of every request and response.

Each document gets an audit record: model family and version, prompt and schema version, raw output, check results, routing decision and reason, reviewer and field-level diff.

evals

Every reviewer correction is a labelled example: the field, the value we extracted and the value the reviewer set. We store them from day one.

From those corrections we build a golden set, stratified by document type and layout, so a layout that is rare in volume still has enough examples to show a regression.

The release gate for any prompt, schema or model change has two parts:

  • field-level accuracy on the golden set must not drop;
  • the auto-accept error rate, measured by sampling auto-accepted documents for review, stays within the owner’s threshold.

Always keep a small random sample of auto-accepted documents going to review. Without it, the auto-accept error rate is invisible, and the golden set only ever sees the documents the checks already doubted.

cost and latency

Cost is driven by OCR pages, model tokens, the number of passes and reviewer minutes. The levers:

  • Skip OCR for born-digital documents.
  • Model tier by document type. A Haiku-class model for simple document types and a Sonnet-class model for complex ones. Use a current model that supports structured outputs.
  • Route by document type early, so each type gets its own prompt, schema and model.
  • Prompt caching of the instructions, which are the same on every call. Keep the schema stable too: its compiled grammar is cached separately, and changing the output format invalidates the prompt cache.
  • Two-pass agreement only on high-consequence fields, not on the whole document.

Batch inference suits large backlogs, but it can’t do the retry turn inline, and AWS documentation disagrees on whether it supports structured outputs. Test it before you plan around it.

To estimate, work from a pilot rather than guesses:

documents per day × (OCR pages + input and output tokens × passes) + review minutes per reviewed document × the review rate.

Measure each term on a sample of real documents, then price it against current published rates for your Region and model. Reviewer minutes are multiplied by the review rate, which is one more reason the threshold decision matters.

operating it

The failure modes we plan runbooks for:

  • OCR or Bedrock throttling. The workflow’s Retry backs off, and a burst drains more slowly rather than failing.
  • Schema rejection after a schema change. An unsupported schema returns a 400, and a new schema’s first compilation can take a few minutes before it is cached, so test every schema change before you deploy it.
  • A growing review backlog. Watch time in queue, and decide in advance who staffs the queue when it grows.
  • A sudden shift in the review rate. Usually a new supplier layout or a model change. Compare against the last release before you touch thresholds.

The dashboards that matter:

  • review rate by document type;
  • time in queue;
  • retry rate;
  • auto-accept sample error rate.

A step change in any of them pages the pipeline’s owner, and a rise in the auto-accept sample error rate also goes to the business owner who set the threshold.

when not to use this

  • Fixed-layout forms that a template extractor already handles well. A model adds cost without adding much.
  • Tiny volumes. If a person can comfortably read every document, build nothing, or a simple assistant.
  • No one will own the threshold or staff the review queue. Without both, the pipeline either reviews everything forever or auto-accepts with no one watching.

on the Claude API directly

The same shape works with the Claude API outside AWS:

  • structured outputs (output_config.format) or strict tool use for the schema;
  • PDFs sent as document blocks, which include page images, so scanned pages can go in directly. The Textract-first step is mainly a constraint of Converse on Bedrock;
  • your own queue and workflow engine in place of Step Functions.

The evidence-based routing and the review loop carry over unchanged. The model fills in the fields, and the evidence decides who checks them.

Related service

This is part of our Build work — make it survive production.