What’s new here
aws-bench provisions real, isolated AWS accounts, deploys CDK stacks into them, and runs the agent with scoped credentials inside a sandboxed container. Scoring for introspection tasks uses an LLM judge; scoring for mutation tasks checks actual AWS resource state programmatically and then rolls the changes back. That gives a reproducible signal on real cloud work.
The project is built on Harbor, an open-source framework for evaluating AI agents and language models, extended with AWS-specific environment provisioning, scenarios, and verifiers.
What it does
aws-bench evaluates AI coding agents (Claude Code, Kiro CLI, and others) on two kinds of AWS tasks:
- Introspection tasks: the agent reads a live environment and diagnoses a misconfiguration; an LLM judge compares the answer to a reference.
- Mutation tasks: the agent creates or modifies real AWS resources; a programmatic verifier checks cloud state and then rolls the changes back.
The benchmark ships three curated datasets: aws-bench-quickstart (9 tasks, 1 scenario), aws-bench-basic (78 tasks, 4 scenarios), and aws-bench-advanced (47 tasks, 3 scenarios). Every scenario is also runnable as a standalone dataset. Task definitions live in the companion repo aws-bench-datasets.
Results land in jobs/<timestamp>/, one reward.json per trial (1.0 = pass, 0.0 = fail).
Who it’s for
Engineers and researchers who need a concrete, reproducible signal on how well an agent handles real AWS operations: diagnosing misconfigurations, provisioning infrastructure, and multi-service troubleshooting. You need an AWS management account with permission to create an AWS Organization, Docker with Compose v2 and buildx ≥ 0.17.0, Python 3.12+, and uv as your package manager.
Try it
Requires macOS or Linux, Python 3.12+, uv, Docker with Compose v2 and buildx ≥ 0.17.0, and an AWS account with permissions to create an Organization and member accounts.
# Install
git clone https://github.com/aws-bench/aws-bench.git && cd aws-bench
uv sync
uv run aws-bench --help
# Configure AWS credentials (static keys)
aws configure --profile my-aws-bench-profile
# ...or IAM Identity Center (SSO):
aws configure sso --profile my-aws-bench-profile # one-time setup
aws sso login --profile my-aws-bench-profile # refresh the session later
export AWS_PROFILE=my-aws-bench-profile
# Set required region
export AWS_REGION=us-east-1
export AWS_DEFAULT_REGION=us-east-1
# Provision environment
uv run aws-bench env init \
--env-name aws-bench-env \
-d aws-bench-quickstart \
--wait-for-quotas
# Deploy scenario resources
uv run aws-bench env setup \
--env-name aws-bench-env \
-d aws-bench-quickstart
# Mint a Bedrock token (if using Bedrock for inference)
eval $(uv run aws-bench env creds --eval)
# Run the benchmark
uv run aws-bench run \
--env-name aws-bench-env \
-d aws-bench-quickstart \
-a claude-code \
-m <model-id> \
--ve AWS_BEARER_TOKEN_BEDROCK=$AWS_BEARER_TOKEN_BEDROCK \
--yesSee the Getting Started guide for detailed setup, account creation, and quota handling.
How mature is it
Created July 2026, with one release (v0.7.0, July 24 2026). Active development: 70 commits in the last 90 days, 17 contributors, 20 forks, 134 stars, 1 open issue and 14 open pull requests. Last pushed October 9 2026. Licensed Apache-2.0. An arXiv paper and a public leaderboard are listed as near-term roadmap items but not yet shipped.