These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.
Weight pin not recorded on this receipt. No download URL.
500 human-verified repository issues. Public-comparable headline for repo software engineering. Agent and harness belong on the cite.
Scoring & limits
swe-bench-verified · 500 tasks
This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.