01 · EVIDENCE
Missing means unknown.
No result is not zero. An empty cell leaves the question open.
The unit of evaluation is a complete setup: model, weight file, quantization, runtime, hardware and harness.
01 · EVIDENCE
No result is not zero. An empty cell leaves the question open.
02 · COMPLETION
Illustration: two successful checks show progress. The job still needs the third.
Local receipts define the benchmarks and exact versions shown here. Published reports provide context within that program; they do not expand it. AutomationBench public and private sets are an explicitly labeled exception. Task sets, protocols, scores and ranking groups stay separate.
Tested by LocalsOnly means we executed the evaluation on the stated setup and retained its results. Published means a vendor or external source reported the result. Published results are references, not independent replications by this lab.
The leaderboard sorts reported scores without assigning a shared rank. Local runs, provider evaluations and comparison tables can use different task sets, tools or budgets. Open a receipt before treating a score gap as a matched comparison. There is no overall score combining these tests.
AutomationBench mean reward measures partial credit. Fully passed tasks are shown separately. A task with some correct steps is not labeled fully passed.
Task failures, timeouts and infrastructure issues belong with the result. Some stored receipts exclude infrastructure errors from their scoring denominator. The SWE-bench Pro receipt, for example, reports 50 reward-present trials and 32 excluded infrastructure errors.
We only compute quality changes when two distinct runs declare the same comparison protocol and benchmark. Current data does not establish a matched quant or hardware comparison. Published-versus-local differences are not isolated quantization losses.
No result is not zero. Memory fit is not a benchmark. Untested hardware, absent latency and missing traces are labeled explicitly. Model specifications are catalog metadata; tested context is taken from the run.