LocalsOnlyevaluationsHow to read this
Method

How to read this

The unit of evaluation is a complete setup: model, weight file, quantization, runtime, hardware and harness.

01 · EVIDENCE

Missing means unknown.

No measurement
0%Measured, no passes

No result is not zero. An empty cell leaves the question open.

02 · COMPLETION

Partial credit is not a completed job.

✓ Check 1✓ Check 2× Check 3→ Incomplete

Illustration: two successful checks show progress. The job still needs the third.

Our local evaluation program

Local receipts define the benchmarks and exact versions shown here. Published reports provide context within that program; they do not expand it. AutomationBench public and private sets are an explicitly labeled exception. Task sets, protocols, scores and ranking groups stay separate.

Two kinds of evidence

Tested by LocalsOnly means we executed the evaluation on the stated setup and retained its results. Published means a vendor or external source reported the result. Published results are references, not independent replications by this lab.

Compare one benchmark at a time

The leaderboard sorts reported scores without assigning a shared rank. Local runs, provider evaluations and comparison tables can use different task sets, tools or budgets. Open a receipt before treating a score gap as a matched comparison. There is no overall score combining these tests.

Partial credit is not a completed job

AutomationBench mean reward measures partial credit. Fully passed tasks are shown separately. A task with some correct steps is not labeled fully passed.

Show the denominator and the failures

Task failures, timeouts and infrastructure issues belong with the result. Some stored receipts exclude infrastructure errors from their scoring denominator. The SWE-bench Pro receipt, for example, reports 50 reward-present trials and 32 excluded infrastructure errors.

A quantization penalty needs a controlled test

We only compute quality changes when two distinct runs declare the same comparison protocol and benchmark. Current data does not establish a matched quant or hardware comparison. Published-versus-local differences are not isolated quantization losses.

Missing means unknown

No result is not zero. Memory fit is not a benchmark. Untested hardware, absent latency and missing traces are labeled explicitly. Model specifications are catalog metadata; tested context is taken from the run.