the problem
Invoices, claims and contracts arrive by the thousand, and today someone keys the fields out of them by hand: supplier, invoice number, date, amount, parties.
The business wants most of those documents processed without a person touching them. But a wrong amount or a wrong party is expensive. So the real requirement isn’t “extract fields”. It is: extract fields, know which ones you can’t vouch for, and send only those to a person.
Around that sit the non-functionals we design for from the start:
- Throughput and bursts. Documents often arrive in batches, at month end for example, not evenly through the day.
- PII. Names, addresses, bank details and ID numbers are often the very fields we’re extracting.
- An audit trail per document. When a figure is questioned, someone will ask what the system read, what it decided and who changed it.
- Cost per document. OCR, model tokens and reviewer minutes all add up.
the design
- A document is uploaded to the inbox bucket in Amazon S3, encrypted with KMS.
- S3 sends an
Object Createdevent to Amazon EventBridge (EventBridge notifications must be enabled on the bucket), and a rule starts a Step Functions workflow. - The workflow’s first step writes an idempotency key to DynamoDB with a conditional write. If the key already exists, this is a duplicate and the execution stops.
- When the document needs it, Textract or Bedrock Data Automation reads it and returns text, layout and per-block confidence. Born-digital, text-only PDFs skip this step.
- Claude on Amazon Bedrock extracts the fields with structured output, constrained to a JSON schema.
- Code checks the result: validators, business rules and grounding in the source. On failure, the errors go back to Claude for one retry, and the result is checked again.
- A
Choicestate routes documents that pass every check to be written to the system of record. - Everything else goes to the review queue. The workflow sends a task token through Amazon SQS and waits on it, a reviewer corrects or confirms the fields in the review UI, and the UI resumes the workflow.
- Every outcome, accepted or reviewed, writes an audit record and the eval data that comes from it.
The review wait uses .waitForTaskToken, which can hold an execution open for up to the one-year Step Functions quota. That is plenty for a queue that empties over a weekend, and it means we use a Standard workflow, not an Express one.
The extraction call uses the Converse outputConfig. Here is a trimmed invoice schema. Every extracted field is nullable, and additionalProperties is false, as the schema subset requires:
{
"type": "object",
"properties": {
"supplier_name": { "type": ["string", "null"] },
"invoice_number": { "type": ["string", "null"] },
"total_amount": { "type": ["string", "null"], "description": "As printed on the document" },
"not_found": {
"type": "array",
"items": {
"type": "object",
"properties": { "field": { "type": "string" }, "reason": { "type": "string" } },
"required": ["field", "reason"],
"additionalProperties": false
}
}
},
"required": ["supplier_name", "invoice_number", "total_amount", "not_found"],
"additionalProperties": false
}
The schema goes into outputConfig as a JSON string:
{
"textFormat": {
"type": "json_schema",
"structure": {
"jsonSchema": {
"schema": "<the schema above, serialised as a string>",
"name": "invoice_extraction"
}
}
}
}
The same schema works as a strict tool’s input schema, with toolChoice set to that tool.
The check step first checks the stop reason, then validates in order: schema, then business rules, then grounding. On failure it sends one follow-up turn with the error list, and then gives up to review:
import json
import boto3
from jsonschema import Draft202012Validator
bedrock = boto3.client("bedrock-runtime")
def extract(model_id, messages, output_config, schema, checks):
"""checks: functions taking the result and returning a list of error strings,
business rules first, then grounding against the document's text."""
validator = Draft202012Validator(schema)
for attempt in range(2): # the first pass plus one retry
resp = bedrock.converse(modelId=model_id, messages=messages, outputConfig=output_config)
if resp["stopReason"] != "end_turn": # refusal or token limit: output may not match
return {"route": "review", "result": None, "errors": [f"stop reason {resp['stopReason']}"]}
reply = resp["output"]["message"]
result = json.loads(reply["content"][0]["text"])
errors = [e.message for e in validator.iter_errors(result)]
if not errors:
errors = [err for check in checks for err in check(result)]
if not errors:
return {"route": "accept", "result": result, "attempts": attempt + 1}
messages = messages + [reply, {"role": "user", "content": [{"text": (
"These values failed our checks. Correct them from the document, "
"or set them to null and list them in not_found:\n" + "\n".join(errors))}]}]
return {"route": "review", "result": result, "errors": errors}
If you use strict tool use instead, the follow-up is a toolResult with status: "error" carrying the same list.
In the workflow, a Choice state reads the route, and anything not accepted goes to a review task that sends the task token to the queue and waits (trimmed):
"Route": {
"Type": "Choice",
"Choices": [
{ "Variable": "$.check.route", "StringEquals": "accept", "Next": "WriteRecord" }
],
"Default": "Review"
},
"Review": {
"Type": "Task",
"Resource": "arn:aws:states:::sqs:sendMessage.waitForTaskToken",
"Parameters": {
"QueueUrl": "https://sqs.<region>.amazonaws.com/<account-id>/extraction-review",
"MessageBody": {
"documentId.$": "$.documentId",
"reason.$": "$.check.errors",
"taskToken.$": "$$.Task.Token"
}
},
"ResultPath": "$.review",
"Next": "WriteRecord"
}
A small consumer turns each message into a row in the task table. When the reviewer submits, the UI calls SendTaskSuccess with that token and the corrected fields, and the workflow carries on with them.
components
| component | AWS service | responsibility |
|---|---|---|
| Inbox | Amazon S3 (KMS-encrypted) | Receives documents, with lifecycle rules on raw files. |
| Event | Amazon EventBridge | Matches Object Created events for the inbox and starts the workflow. |
| Dedupe | Amazon DynamoDB | Holds idempotency keys, written with a conditional write so each document is processed once. |
| Workflow | AWS Step Functions | Orchestrates read, extract, check and route, and owns throttling retries. |
| Read | Amazon Textract or Bedrock Data Automation | OCR and layout for scanned or layout-heavy documents, with per-block confidence. |
| Extract | Claude on Amazon Bedrock | Fills the schema from the document’s text and layout. Has no tools with side effects. |
| Check | your code (Lambda or a workflow step) | Validators, business rules and grounding. Produces the routing decision and its reason. |
| Route | Step Functions Choice state | Sends each document to accept or to review, based on the check results. |
| Accept | your system of record | Receives the accepted fields. |
| Review queue | a queue we own: Amazon SQS plus a DynamoDB task table | The workflow sends each waiting document’s task token and review reason to SQS; a small consumer writes it to the task table the review UI reads. |
| Reviewer | your review UI | Shows the document and the flagged fields, and calls SendTaskSuccess with the corrected values. |
| Audit and eval data | Amazon S3 and DynamoDB | One record per document, plus reviewer corrections as labelled examples. |
Amazon A2I is in maintenance mode and closed to new customers since July 2026, so we don’t build new pipelines on it. The queue we own is a small amount of code, and it keeps the review data where the evals can use it.
decisions
How Claude sees the document
On Converse, a PDF document block is read as text only unless citations are on, and citations can’t be combined with structured outputs. With the schema we want, a scanned page or a table’s layout would be lost.
Running OCR and layout extraction first gives Claude the text and its layout, plus a per-block confidence from Textract that the grounding check uses later. For born-digital, text-only documents, skip this step and send the PDF directly.
Page images and InvokeModel with a document block both give Claude the visual page, and are worth testing for layout-heavy documents. The trade-off of our choice is one more service and one more cost line.
How the output shape is enforced
Constrained decoding guarantees the result parses against the schema on a normal stop, so no more repair prompts for broken JSON. A refusal or a stop at the token limit can still return output that doesn’t match, so the code checks the stop reason first.
Give each field type: ["string", "null"] so the model can say “not found” instead of inventing a value, and add a not_found list for the reason. The schema subset can’t express ranges or lengths, so those checks live in code. Strict tool use is equivalent if the extraction is one step of a larger tool-using flow.
The schema guarantees the shape of the answer, not that it is true.
Where the routing signal comes from
We route on four kinds of evidence:
- deterministic validators, such as formats and check digits;
- business rules: totals add up, dates are in range, the supplier exists;
- a grounding check that each value appears in the source text, or carries a Textract or Bedrock Data Automation confidence;
- agreement between two passes, for high-consequence fields only.
A required field that comes back null is a signal too. A confidence number the model reports about itself is at most a tie-breaker, never a gate.
How many retries before review
One follow-up turn that feeds back the specific validation errors fixes most slips. Beyond that, retries cost money and hide real ambiguity, which is exactly what a person should see.
Throttling and timeouts are a different kind of failure. They are handled by the Step Functions Retry with backoff, not by this budget.
Who sets the review threshold
How much error is acceptable is a business decision. Amount fields and party names deserve stricter rules than a reference field.
Start with every document reviewed. Then let auto-accept earn its way in, field by field, as review outcomes show the checks are reliable. Name the owner, and write their thresholds down next to the rules they govern.
How duplicates are stopped
S3 event notifications are at-least-once, and duplicates carry the same key, version ID and sequencer.
Key on bucket, key and version ID, or on a SHA-256 of the content, and record it with a DynamoDB conditional write. Where you start executions yourself, use the same key as the Step Functions execution name, hashed to fit the naming rules, so Step Functions won’t start a second execution for the same document.
Separately, catch the same invoice arriving twice under different file names, for example the same supplier and invoice number. That check belongs with the business rules, and a hit goes to review.
safety and governance
The document is untrusted input. Anyone who can send an invoice can put text in it that reads like instructions. So:
- it goes in only as document or text content with a neutral name, never as instructions, and never in the system prompt;
- the extraction step has no tools with side effects, so the worst a malicious document can do is plant a wrong value;
- validators and grounding are there to catch planted values.
For PII:
- KMS on the bucket, and PrivateLink to Bedrock so traffic stays off the public internet;
- tag PII fields with Comprehend
DetectPiiEntities, or Guardrails in Detect mode, but don’t mask before extraction when the PII is the field; - mask PII in logs and in the reviewer UI unless the reviewer’s role needs it;
- lifecycle rules on raw documents, matched to your retention policy;
- restrict model-invocation logging, which, when enabled, can store the full text of every request and response.
Each document gets an audit record: model family and version, prompt and schema version, raw output, check results, routing decision and reason, reviewer and field-level diff.
evals
Every reviewer correction is a labelled example: the field, the value we extracted and the value the reviewer set. We store them from day one.
From those corrections we build a golden set, stratified by document type and layout, so a layout that is rare in volume still has enough examples to show a regression.
The release gate for any prompt, schema or model change has two parts:
- field-level accuracy on the golden set must not drop;
- the auto-accept error rate, measured by sampling auto-accepted documents for review, stays within the owner’s threshold.
Always keep a small random sample of auto-accepted documents going to review. Without it, the auto-accept error rate is invisible, and the golden set only ever sees the documents the checks already doubted.
cost and latency
Cost is driven by OCR pages, model tokens, the number of passes and reviewer minutes. The levers:
- Skip OCR for born-digital documents.
- Model tier by document type. A Haiku-class model for simple document types and a Sonnet-class model for complex ones. Use a current model that supports structured outputs.
- Route by document type early, so each type gets its own prompt, schema and model.
- Prompt caching of the instructions, which are the same on every call. Keep the schema stable too: its compiled grammar is cached separately, and changing the output format invalidates the prompt cache.
- Two-pass agreement only on high-consequence fields, not on the whole document.
Batch inference suits large backlogs, but it can’t do the retry turn inline, and AWS documentation disagrees on whether it supports structured outputs. Test it before you plan around it.
To estimate, work from a pilot rather than guesses:
documents per day × (OCR pages + input and output tokens × passes) + review minutes per reviewed document × the review rate.
Measure each term on a sample of real documents, then price it against current published rates for your Region and model. Reviewer minutes are multiplied by the review rate, which is one more reason the threshold decision matters.
operating it
The failure modes we plan runbooks for:
- OCR or Bedrock throttling. The workflow’s
Retrybacks off, and a burst drains more slowly rather than failing. - Schema rejection after a schema change. An unsupported schema returns a 400, and a new schema’s first compilation can take a few minutes before it is cached, so test every schema change before you deploy it.
- A growing review backlog. Watch time in queue, and decide in advance who staffs the queue when it grows.
- A sudden shift in the review rate. Usually a new supplier layout or a model change. Compare against the last release before you touch thresholds.
The dashboards that matter:
- review rate by document type;
- time in queue;
- retry rate;
- auto-accept sample error rate.
A step change in any of them pages the pipeline’s owner, and a rise in the auto-accept sample error rate also goes to the business owner who set the threshold.
when not to use this
- Fixed-layout forms that a template extractor already handles well. A model adds cost without adding much.
- Tiny volumes. If a person can comfortably read every document, build nothing, or a simple assistant.
- No one will own the threshold or staff the review queue. Without both, the pipeline either reviews everything forever or auto-accepts with no one watching.
on the Claude API directly
The same shape works with the Claude API outside AWS:
- structured outputs (
output_config.format) or strict tool use for the schema; - PDFs sent as document blocks, which include page images, so scanned pages can go in directly. The Textract-first step is mainly a constraint of Converse on Bedrock;
- your own queue and workflow engine in place of Step Functions.
The evidence-based routing and the review loop carry over unchanged. The model fills in the fields, and the evidence decides who checks them.