Data systems

aws-samples/

sample-synthetic-document-generator

CLI tool that uses Amazon Bedrock to rewrite a reference PDF as a synthetic copy with fake data, then audits the output for PII with Amazon Comprehend.

What’s new here

It starts from a real PDF: a paid Bedrock call extracts structure and field observations, then generate produces thousands of synthetic rows for free, offline, indefinitely; documents still need the paid convert. The cost gate is explicit: estimate before you commit, then verify after you generate to scan the rows for leaked source PII.

What it does

Give pocsynth convert a PDF and it produces a synthetic HTML document with the same layout, section structure, and table shapes, but with all prose rewritten and names, addresses, and identifiers replaced by realistic fakes written by the model. Amazon Comprehend then scans the output and writes a PII audit CSV.

Pass --num-docs N to get N independent rewrites from one source. Each lands in its own directory with per-page HTML and PNG files.

Beyond document conversion, a second pipeline handles tabular data. extract pulls field observations from a PDF, schema turns them into a generation-ready schema, and generate produces typed, seeded CSV or JSON rows offline, with PII fields bound to Faker providers. Fifteen bundled presets (support tickets, insurance claims, financial transactions, and more) let you skip the extraction step at zero cost.

A web UI ships as an optional extra (uv tool install '.[ui]'). Built with FastAPI and HTMX, it lets you compose a dataset from record-type pills, preview 10 rows, and download any size. Every preview also shows the equivalent CLI command and an agent-skill invocation.

Who it’s for

Solutions Architects and PoC engineers who need shape-correct documents for demos, RAG evaluation corpora, or partner handoffs without sharing real contracts or forms. The structured-data pipeline also suits anyone who needs reproducible, typed synthetic rows from a reference document, with only one or two paid Bedrock calls.

Try it

# Install
git clone https://github.com/aws-samples/sample-synthetic-document-generator.git
cd sample-synthetic-document-generator
uv tool install .
 
# Verify AWS wiring
pocsynth doctor
 
# Grab the sample PDF
curl -sSfLo aws-mp-contract.pdf \
  https://s3.amazonaws.com/aws-mp-standard-contracts/Standard-Contact-for-AWS-Marketplace-2022-07-14.pdf
 
# Pre-flight cost estimate
pocsynth estimate aws-mp-contract.pdf
 
# Convert with PII audit
pocsynth convert aws-mp-contract.pdf --redact-values
 
# Generate synthetic rows from a bundled preset (free, no AWS needed)
pocsynth generate --preset crm_contacts --rows 1000 --seed 42 -o ./out
 
# Optional web UI
uv tool install '.[ui]'
pocsynth ui   # http://127.0.0.1:8000

Requires Python 3.10+, uv on PATH, and AWS credentials with bedrock:Converse and comprehend:DetectPiiEntities. The README includes a minimum IAM policy.

How mature is it

Created May 2026, last pushed October 2026. 32 stars, 0 forks, 2 contributors, 1 commit in the last 90 days, 4 open issues and pull requests, no releases. Licensed MIT-0. The README lists 481 stubbed unit tests and 14 live tests. This is an AWS Samples PoC accelerator, not a production service.