Agent evaluation

fstandhartinger/

jevbench

Benchmark and leaderboard for AI decision models, scoring typed accuracy, probability calibration, latency, and cost per 1,000 decisions.

What’s new here

JevBench measures a specific task: given a state, a rubric, and a fixed label set, does the model pick the right label, and does its stated probability distribution mean anything? Calibration (ECE and fidelity to exact gold distributions) is a scored axis with equal weight. Cost is reported per 1,000 decisions, not per 1,000 tokens, which matters because a single decision can run to hundreds or thousands of input tokens.

What it does

JevBench runs a suite of decision tasks across six families: routing, answer adequacy judging, policy yes/no checks, intent classification, ordinal severity scoring, and enum extraction. Each system sees identical state, instructions, rubric, and label set; only the transport adapter differs.

Four axes are scored 0-100: Intelligence (chance-corrected accuracy, with a growing penalty below 50), Calibration (ECE and label-distribution fidelity on the hard tier), Speed (log-scale latency at p50 and p95), and Cost (log-scale dollars per 1,000 decisions at the provider’s published tariff). The Capability Score is the arithmetic mean of Intelligence and Calibration. The Composite Score is a weighted harmonic mean of all four axes.

The historical v1.4.2.2 board has 95 systems, 91 ranked. The current published release is v1.6.1 (live at benchmarkheaven.com/jev-models). Historical releases keep their original scoring definitions; scores across method versions are not comparable.

Who it’s for

Teams building or evaluating decision models that must pick from a typed label set. Also useful for anyone who wants to compare hosted decision APIs on calibration and cost rather than just accuracy. The harness is also usable locally to run the public 72-decision cohort against any Jev-compatible endpoint.

Try it

Python 3.10+, run from a clone of the repo. HTTP adapters need only the standard library. The Jev run needs a TypeSafe API key in TYPESAFE_API_KEY; the open-rebuild run needs a public endpoint URL in place of the placeholder.

python -m unittest discover -s tests -v
 
# Jev, the published 72-decision cohort
python -m jevbench.cli run --tasks datasets/public/original.jsonl \
  --adapter typesafe --model jev-latest --key-env TYPESAFE_API_KEY \
  --price-in-per-m 0.042 --price-out-per-m 0 \
  --results RUN/results.jsonl --raw-dir RUN/raw \
  --ledger RUN/ledger.jsonl --cap-usd 15 --manifest RUN/manifest.json
 
# an open rebuild on its author's public endpoint - note the empty key
python -m jevbench.cli run --tasks datasets/public/original.jsonl \
  --adapter typesafe --endpoint https://SOME-PUBLIC-ENDPOINT --key-env '' \
  --model jev-latest --cost-basis no_billable_account_public_endpoint \
  --reserve-usd 0 --delay-s 0.2 \
  --results RUN2/results.jsonl --raw-dir RUN2/raw --ledger RUN/ledger.jsonl
 
python -m jevbench.cli summarize --tasks datasets/public/original.jsonl \
  --results RUN/results.jsonl --public-export RUN/summary.json

The harness uses a file-locked cost ledger: it reserves worst-case cost before each request and settles after. A crashed run’s reservation stays charged. Every run directory is created fresh; results are never overwritten.

How mature is it

251 stars, 25 forks, 3 contributors. Created 2026-09-19; latest push 2026-10-10. 81 commits in the last 90 days. 3 releases; latest tagged v1.4.2 (2026-09-24), with the live board at v1.6.1 (2026-10-06). 158 open issues and 8 open pull requests. MIT license for the harness and the 72 original public decisions.