What’s new here
Commerce Agent Bench tests whether an agent can complete multi-step business workflows: publishing a product catalog entry, booking freight, customizing a storefront theme. Each task runs in a fresh Docker container against local mock replicas of real commerce platforms, and passes only when every verifier check clears against actual state changes, not just model output.
What it does
The v1.3.1 suite has 107 tasks split across four interface types: 53 CLI, 28 browser, 16 file, and 10 API/MCP. Tasks cover supplier analysis, product publishing, logistics, document production, public-web research, and commerce operations. Three capability slices let you run text-only, browser-plus-text, or vision-required subsets.
Each run spins up a container, stages only the agent-visible task tree (task instructions and a workspace), and keeps graders, rubrics, and mock service state out of agent reach. After the agent exits, a host-side verifier checks outputs and mock state, writes a reward record, and archives the full trajectory, logs, screenshots, and container metadata.
Reference scores for 13 model families are published across three harnesses: Pi, OpenClaw, and Accio. Claude Opus 5 leads all three, topping out at 61.7% on the Accio harness. The benchmark pins all four reproducibility components (task set, task definitions, harness, and runtime image digest) so results from runs using the same benchmark version can be compared, though the README notes that routing, model snapshots, prompt adapters, retry policies, and judge endpoints all affect outcomes.
Who it’s for
AI researchers and model teams who want to evaluate agents on commerce and business-process workflows. Also useful for teams building commerce agents who want a graded offline testbed before touching production systems. The Accio team at Alibaba International will run evaluations on request, including private pre-release checkpoints.
Try it
Requires Docker with Linux container support, Python 3.11 or newer, a model API key, and an LLM-judge API key.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
commerce-agent-bench listPull the pinned runtime image:
docker pull --platform linux/amd64 \
acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859Run a single task with the native Gemini route:
export GEMINI_API_KEY="..."
commerce-agent-bench run api-amazon-margin-floor-audit \
--harness openclaw \
--image acciolyk/accio_bench@sha256:1e9cf5c72a56794175b7d06ece036b92e296e6b7e9e9a7fa244026f6acea3859 \
--platform linux/amd64 \
--openclaw-model google/gemini-3.5-flash \
--openclaw-image-model google/gemini-3.5-flash \
--openclaw-models-config configs/native_google_direct_models.json \
--llm-judge-provider gemini \
--llm-judge-model gemini-3.1-pro-preview \
--run-id smokeFor a full collection run: commerce-agent-bench run --config configs/openclaw_native_google_direct.yaml --run-id "openclaw-$(date +%Y%m%d-%H%M%S)". Use --limit 1 for a quick smoke test before committing to the full suite.
How mature is it
1,278 stars, 70 forks, 2 contributors, 33 commits in the past 90 days, 0 formal releases (versioned via badges at v1.3.1). Created August 2026, last pushed August 2026. Harness and Python package are Apache-2.0; the task suite under datasets_domain_v1/ is CC BY 4.0.