Theme

Agent evaluation

5 repos reviewed · since Oct 10, 2026 · updated Oct 10, 2026 · rewritten each time a repo joins

In one paragraph

Agent evaluation covers tools that give AI agents or models a defined job and check whether they did it correctly. aws-bench spins up real AWS accounts and checks live cloud state. CommerceAgentBench runs 107 commerce workflows against local mock services. eval holds the model and tasks constant to compare nine coding-agent harnesses. jevbench scores decision models on accuracy, calibration, latency, and cost. SwarmWorld measures collective behavior across societies of agents. What counts here: benchmarks, leaderboards, and research environments that measure how well agents or models do defined work.

The main approaches

Harness benchmarks

Run the same tasks through different agent runtimes to isolate how much the harness itself affects outcomes. frontier-harness-eval/eval holds model, tasks, and vCPU constant across nine harnesses and reports pass rate, cost, cache behavior, and speed.

Domain task benchmarks

Give agents realistic work in a specific domain, then score it with verifiers that check actual system state where a task changes it. aws-bench/aws-bench uses real AWS accounts; Accio-org/CommerceAgentBench uses stateful local mock services for commerce workflows.

Decision-model scoring

Score models specifically on classification and routing decisions, not on freeform generation. fstandhartinger/jevbench measures typed accuracy, calibration, latency, and cost per 1,000 decisions across six decision-task families.

Research environments

Provide a controllable simulation where agents act over many steps and results can be replayed or analyzed. lamm-mit/SwarmWorld runs societies of LLM agents in a shared 2-D world with deterministic replay and counterfactual analysis.

Map of the theme

flowchart LR
  t["Agent evaluation"]
  t --> f1["Harness benchmarks"]
  t --> f2["Domain task benchmarks"]
  t --> f3["Decision-model scoring"]
  t --> f4["Research environments"]
  f1 --> r1["frontier-harness-eval/eval"]
  f2 --> r2["aws-bench/aws-bench"]
  f2 --> r3["Accio-org/CommerceAgentBench"]
  f3 --> r4["fstandhartinger/jevbench"]
  f4 --> r5["lamm-mit/SwarmWorld"]

Where the new ideas are

  • frontier-harness-eval/eval isolates harness-level effects by keeping everything else constant, finding 17.5x cost differences at similar pass rates across nine harnesses.
  • fstandhartinger/jevbench uses a cost metric denominated per 1,000 decisions rather than per 1,000 tokens, and scores a calibration axis (ECE and label-distribution fidelity) alongside accuracy.
  • lamm-mit/SwarmWorld measures collective technological evolution: artifacts persist between ticks, programs are inherited by successors, and deterministic replay enables contributor-removal counterfactuals.
  • aws-bench/aws-bench uses disposable real AWS accounts rather than mocks, so mutation tasks check actual cloud resource state and roll back changes afterward.
  • Accio-org/CommerceAgentBench stages the agent workspace to hide graders and rubrics, then checks both file outputs and mock service state from the host side after the agent exits.

Side by side

RepoApproachTask countReal or mock environmentCalibration scoringReplay / counterfactual
aws-bench/aws-benchDomain task benchmarks134 tasks across 3 datasetsReal AWS accountsNot statedNot stated
Accio-org/CommerceAgentBenchDomain task benchmarks107 tasksLocal mock servicesNot statedNot stated
frontier-harness-eval/evalHarness benchmarks30 tasksNot statedNot statedNot stated
lamm-mit/SwarmWorldResearch environmentsNot stated2-D simulation worldNot statedDeterministic replay with counterfactual contributor-removal
fstandhartinger/jevbenchDecision-model scoring6 task families; a 72-decision public cohortNot statedYes (ECE and label-distribution fidelity)Not stated

How the idea moved

flowchart LR
  n1["started Jul 2026<br/>aws-bench/aws-bench"]
  n2["started Aug 2026<br/>Accio-org/CommerceAgentBench"]
  n3["started Aug 2026<br/>frontier-harness-eval/eval"]
  n4["started Sep 2026<br/>lamm-mit/SwarmWorld"]
  n5["started Sep 2026<br/>fstandhartinger/jevbench"]
  n1 --> n2 --> n3 --> n4 --> n5
  • started Jul 2026 · aws-bench · Adds: Evaluates agents on real AWS infrastructure by provisioning disposable accounts, running agents with scoped credentials, and scoring introspection tasks with an LLM judge and mutation tasks against live cloud resource state.
  • started Aug 2026 · CommerceAgentBench · Adds: Provides 107 long-horizon commerce tasks across CLI, browser, file, and API interfaces, verified against stateful local mock services with graders kept out of agent reach.
  • started Aug 2026 · eval · Adds: Holds model, tasks, and hardware constant across nine coding-agent harnesses to isolate harness-level effects, exposing a 17.5x cost spread at similar pass rates and shipping a reusable evaluation skill for third-party harnesses.
  • started Sep 2026 · SwarmWorld · Adds: Simulates societies of LLM agents in a deterministic 2-D world where artifacts persist across ticks and counterfactual contributor-removal analysis measures collective technological evolution.
  • started Sep 2026 · jevbench · Adds: Scores decision models on typed accuracy, probability calibration (ECE), latency, and cost per 1,000 decisions across six decision-task families, with a composite leaderboard.

Easily confused

This theme could be confused with general LLM benchmarks (MMLU, HumanEval) that test a model’s knowledge or coding ability in isolation. The difference is scope: aws-bench, CommerceAgentBench, frontier-harness-eval and SwarmWorld put a model inside a runtime loop with tools, state, and a real or simulated environment, then check whether the agent completed a workflow, not just whether it answered a question correctly. jevbench scores single typed decisions. It can also look like agent frameworks or sandboxes, but those build or run agents; evaluation measures them.

Gaps nobody has filled

  • Nothing here measures multi-turn collaboration between heterogeneous agents (different models or roles) on a shared task.
  • No benchmark here tests agent behavior against adversarial inputs, such as misleading tool outputs.
  • Cost and speed are scored in jevbench and frontier-harness-eval, but aws-bench lists no cost or latency metric and CommerceAgentBench reports tokens and time only as descriptive telemetry, so efficiency cannot be compared across these suites.