In one paragraph
Agent evaluation covers tools that give AI agents or models a defined job and check whether they did it correctly. aws-bench spins up real AWS accounts and checks live cloud state. CommerceAgentBench runs 107 commerce workflows against local mock services. eval holds the model and tasks constant to compare nine coding-agent harnesses. jevbench scores decision models on accuracy, calibration, latency, and cost. SwarmWorld measures collective behavior across societies of agents. What counts here: benchmarks, leaderboards, and research environments that measure how well agents or models do defined work.
The main approaches
Harness benchmarks
Run the same tasks through different agent runtimes to isolate how much the harness itself affects outcomes. frontier-harness-eval/eval holds model, tasks, and vCPU constant across nine harnesses and reports pass rate, cost, cache behavior, and speed.
Domain task benchmarks
Give agents realistic work in a specific domain, then score it with verifiers that check actual system state where a task changes it. aws-bench/aws-bench uses real AWS accounts; Accio-org/CommerceAgentBench uses stateful local mock services for commerce workflows.
Decision-model scoring
Score models specifically on classification and routing decisions, not on freeform generation. fstandhartinger/jevbench measures typed accuracy, calibration, latency, and cost per 1,000 decisions across six decision-task families.
Research environments
Provide a controllable simulation where agents act over many steps and results can be replayed or analyzed. lamm-mit/SwarmWorld runs societies of LLM agents in a shared 2-D world with deterministic replay and counterfactual analysis.
Map of the theme
flowchart LR t["Agent evaluation"] t --> f1["Harness benchmarks"] t --> f2["Domain task benchmarks"] t --> f3["Decision-model scoring"] t --> f4["Research environments"] f1 --> r1["frontier-harness-eval/eval"] f2 --> r2["aws-bench/aws-bench"] f2 --> r3["Accio-org/CommerceAgentBench"] f3 --> r4["fstandhartinger/jevbench"] f4 --> r5["lamm-mit/SwarmWorld"]
Where the new ideas are
- frontier-harness-eval/eval isolates harness-level effects by keeping everything else constant, finding 17.5x cost differences at similar pass rates across nine harnesses.
- fstandhartinger/jevbench uses a cost metric denominated per 1,000 decisions rather than per 1,000 tokens, and scores a calibration axis (ECE and label-distribution fidelity) alongside accuracy.
- lamm-mit/SwarmWorld measures collective technological evolution: artifacts persist between ticks, programs are inherited by successors, and deterministic replay enables contributor-removal counterfactuals.
- aws-bench/aws-bench uses disposable real AWS accounts rather than mocks, so mutation tasks check actual cloud resource state and roll back changes afterward.
- Accio-org/CommerceAgentBench stages the agent workspace to hide graders and rubrics, then checks both file outputs and mock service state from the host side after the agent exits.
Side by side
| Repo | Approach | Task count | Real or mock environment | Calibration scoring | Replay / counterfactual |
|---|---|---|---|---|---|
| aws-bench/aws-bench | Domain task benchmarks | 134 tasks across 3 datasets | Real AWS accounts | Not stated | Not stated |
| Accio-org/CommerceAgentBench | Domain task benchmarks | 107 tasks | Local mock services | Not stated | Not stated |
| frontier-harness-eval/eval | Harness benchmarks | 30 tasks | Not stated | Not stated | Not stated |
| lamm-mit/SwarmWorld | Research environments | Not stated | 2-D simulation world | Not stated | Deterministic replay with counterfactual contributor-removal |
| fstandhartinger/jevbench | Decision-model scoring | 6 task families; a 72-decision public cohort | Not stated | Yes (ECE and label-distribution fidelity) | Not stated |
How the idea moved
flowchart LR n1["started Jul 2026<br/>aws-bench/aws-bench"] n2["started Aug 2026<br/>Accio-org/CommerceAgentBench"] n3["started Aug 2026<br/>frontier-harness-eval/eval"] n4["started Sep 2026<br/>lamm-mit/SwarmWorld"] n5["started Sep 2026<br/>fstandhartinger/jevbench"] n1 --> n2 --> n3 --> n4 --> n5
- started Jul 2026 · aws-bench · Adds: Evaluates agents on real AWS infrastructure by provisioning disposable accounts, running agents with scoped credentials, and scoring introspection tasks with an LLM judge and mutation tasks against live cloud resource state.
- started Aug 2026 · CommerceAgentBench · Adds: Provides 107 long-horizon commerce tasks across CLI, browser, file, and API interfaces, verified against stateful local mock services with graders kept out of agent reach.
- started Aug 2026 · eval · Adds: Holds model, tasks, and hardware constant across nine coding-agent harnesses to isolate harness-level effects, exposing a 17.5x cost spread at similar pass rates and shipping a reusable evaluation skill for third-party harnesses.
- started Sep 2026 · SwarmWorld · Adds: Simulates societies of LLM agents in a deterministic 2-D world where artifacts persist across ticks and counterfactual contributor-removal analysis measures collective technological evolution.
- started Sep 2026 · jevbench · Adds: Scores decision models on typed accuracy, probability calibration (ECE), latency, and cost per 1,000 decisions across six decision-task families, with a composite leaderboard.
Easily confused
This theme could be confused with general LLM benchmarks (MMLU, HumanEval) that test a model’s knowledge or coding ability in isolation. The difference is scope: aws-bench, CommerceAgentBench, frontier-harness-eval and SwarmWorld put a model inside a runtime loop with tools, state, and a real or simulated environment, then check whether the agent completed a workflow, not just whether it answered a question correctly. jevbench scores single typed decisions. It can also look like agent frameworks or sandboxes, but those build or run agents; evaluation measures them.
Gaps nobody has filled
- Nothing here measures multi-turn collaboration between heterogeneous agents (different models or roles) on a shared task.
- No benchmark here tests agent behavior against adversarial inputs, such as misleading tool outputs.
- Cost and speed are scored in jevbench and frontier-harness-eval, but aws-bench lists no cost or latency metric and CommerceAgentBench reports tokens and time only as descriptive telemetry, so efficiency cannot be compared across these suites.