Tests software maintenance: finding the relevant code and producing a change that resolves an issue.
01Repository + issue→
02Inspect and change code→
03Repository tests
Interpret the score Our local result covers a 50-task subset. It is not interchangeable with a full published evaluation.
BEYOND THE HEADLINE
What happened across 50 tasks?
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Fully passed10 20.0%
No credit40 80.0%
Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.
The failure analysis distinguishes incorrect solutions, syntax errors, wrong files, instruction following, tool use, context limits and loops. The diagnosis is richer than a single repository-fix score.
The LocalsOnly measured receipt records each task as pass or fail.
32 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.
Published rows are vendor or bench cites. They do not share the measured task grid.
SuiteSWE-bench Pro
↓
Tasks50 on the measured receipt
↓
PassRecorded pass on the task
FailRecorded fail on the task
Setup issueExecution failure on the receipt
Follow one task · this measured recipe
A passed task and a failed one
Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-03. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.
scale-ai/instance_protonmail__webclients-32ff10999a06455cb2147f6873d627456924ae13 · Task ID scale-ai/instance_protonmail__webclients-32ff10999a06455cb2147f6873d627456924ae13
Did not complete the recorded checks. The record does not establish a cause.
Exact task metadata
scale-ai/instance_flipt-io__flipt-a0cbc0cb65ae601270bdbe3f5313e2dfd49c80e4 · Task ID scale-ai/instance_flipt-io__flipt-a0cbc0cb65ae601270bdbe3f5313e2dfd49c80e4
This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking
Results & published references
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.