Agent evaluation

frontier-harness-eval/

eval

Benchmark that ran 9 coding-agent harnesses on 30 tasks with one model, exposing 17.5x cost differences at similar pass rates.

What’s new here

FrontierHarness Eval fixes one model (Kimi K3, served by Fireworks), 30 tasks, nine harnesses, and 12 configurations for 360 runs total. The result: harness choice moves cost from $1.05 to $18.34 per passing task while pass rates stay within a ~17-point band (50% to 66.7%). The repo also ships a reproducible skill so a third party can run the same task set against a new harness. A matched control run under the same policy and environment is needed before claiming comparability, and new runs default to receiving no leaderboard rank.

What it does

The repository holds three things: the benchmark definition (benchmark.json), the published results (results/eval-data.json), and an agent-neutral evaluation skill (skills/frontierharness-eval/).

The skill drives a full run: it freezes a golden checkpoint of the harness under test, restores that checkpoint fresh for each task, runs every trial, scores pass/fail with deterministic verifiers, and builds a cost and speed report. Resuming a partial run reuses the same run ID; changing the egress policy requires a new one.

Comparability rules are explicit: use Kimi K3, one fresh restore per task, identical vCPU and memory, infrastructure failures marked infra_invalid not scored as failures. Any relaxed invariant is recorded in the report.

Who it’s for

Teams building or shipping a coding-agent harness who want to know where they sit on pass rate and cost against a fixed reference set. Also researchers who want task definitions and raw result data for 360 harness-task pairs without rerunning anything.

Try it

Install the evaluation skill in the project where you use your coding agent (Node.js and Git required):

npx skills add https://runta.com/docs --skill runta-installer runta-cli
npx skills add frontier-harness-eval/eval --skill frontierharness-eval

Then start a new agent session and ask it to run a quick smoke test:

Use the frontierharness-eval skill to evaluate https://github.com/acme/my-harness.
Start with one Terminal-Bench task and one DeepSWE task.

To query the published results directly, from a clone of the repo:

jq '.harnesses[] | {name, successful, effective_cost_per_pass}' results/eval-data.json

How mature is it

Created August 2026, last pushed September 8, 2026. 304 stars, 21 forks, 2 contributors, 44 commits in the past 90 days, 11 open issues and 8 open pull requests. No releases and no license. Benchmark is labeled v1.0. Sponsored by Runta, which provided the isolated runtimes for all 360 baseline evaluations.