Agent evaluation

Accio-org/

CommerceAgentBench

107-task benchmark for long-horizon commerce agents: product publishing, freight booking, storefront ops, all in stateful local mock services.

What’s new here

Commerce Agent Bench tests whether an agent can complete multi-step business workflows: publishing a product catalog entry, booking freight, customizing a storefront theme. Each task runs in a fresh Docker container against local mock replicas of real commerce platforms, and passes only when every verifier check clears against actual state changes, not just model output.

What it does

The v1.3.1 suite has 107 tasks split across four interface types: 53 CLI, 28 browser, 16 file, and 10 API/MCP. Tasks cover supplier analysis, product publishing, logistics, document production, public-web research, and commerce operations. Three capability slices let you run text-only, browser-plus-text, or vision-required subsets.

Each run spins up a container, stages only the agent-visible task tree (task instructions and a workspace), and keeps graders, rubrics, and mock service state out of agent reach. After the agent exits, a host-side verifier checks outputs and mock state, writes a reward record, and archives the full trajectory, logs, screenshots, and container metadata.

Reference scores for 13 model families are published across three harnesses: Pi, OpenClaw, and Accio. Claude Opus 5 leads all three, topping out at 61.7% on the Accio harness. The benchmark pins all four reproducibility components (task set, task definitions, harness, and runtime image digest) so results from runs using the same benchmark version can be compared, though the README notes that routing, model snapshots, prompt adapters, retry policies, and judge endpoints all affect outcomes.

Who it’s for

AI researchers and model teams who want to evaluate agents on commerce and business-process workflows. Also useful for teams building commerce agents who want a graded offline testbed before touching production systems. The Accio team at Alibaba International will run evaluations on request, including private pre-release checkpoints.

Try it

Requires Docker with Linux container support, Python 3.11 or newer, a model API key, and an LLM-judge API key.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
commerce-agent-bench list

Pull the pinned runtime image:

docker pull --platform linux/amd64 \
  acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859

Run a single task with the native Gemini route:

export GEMINI_API_KEY="..."
 
commerce-agent-bench run api-amazon-margin-floor-audit \
  --harness openclaw \
  --image acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859 \
  --platform linux/amd64 \
  --openclaw-model google/gemini-3.5-flash \
  --openclaw-image-model google/gemini-3.5-flash \
  --openclaw-models-config configs/native_google_direct_models.json \
  --llm-judge-provider gemini \
  --llm-judge-model gemini-3.1-pro-preview \
  --run-id smoke

For a full collection run: commerce-agent-bench run --config configs/openclaw_native_google_direct.yaml --run-id "openclaw-$(date +%Y%m%d-%H%M%S)". Use --limit 1 for a quick smoke test before committing to the full suite.

How mature is it

1,278 stars, 70 forks, 2 contributors, 33 commits in the past 90 days, 0 formal releases (versioned via badges at v1.3.1). Created August 2026, last pushed August 2026. Harness and Python package are Apache-2.0; the task suite under datasets_domain_v1/ is CC BY 4.0.