LocalsOnlyevaluationsHow to read this
← Tests

CodingSoftware engineering

SWE-bench Pro

Pro evaluation · 50 tasks. Exact task sets and protocols remain separate.

Can it fix an issue in an existing codebase?

20.0%Recorded score · 10 / 50 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Can it fix an issue in an existing codebase?

Tests software maintenance: finding the relevant code and producing a change that resolves an issue.

  1. 01Repository + issue
  2. 02Inspect and change code
  3. 03Repository tests

Interpret the score Our local result covers a 50-task subset. It is not interchangeable with a full published evaluation.

BEYOND THE HEADLINE

What happened across 50 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed10 20.0%
No credit40 80.0%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
What this test is

What this benchmark measures

Resolve issues in real software repositories.

The failure analysis distinguishes incorrect solutions, syntax errors, wrong files, instruction following, tool use, context limits and loops. The diagnosis is richer than a single repository-fix score.

SWE-bench Pro paper
How it grades

How a task earns its score

The LocalsOnly measured receipt records each task as pass or fail.

32 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.

Published rows are vendor or bench cites. They do not share the measured task grid.

SuiteSWE-bench Pro
Tasks50 on the measured receipt
PassRecorded pass on the task
FailRecorded fail on the task
Setup issueExecution failure on the receipt
Follow one task · this measured recipe

A passed task and a failed one

Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-03. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.

passed vs not passed

PassedPassed on this receipt

Scale ai/instance protonmail webclients 32ff10999a06455cb2147f6873d627456924ae13

Completed the recorded checks.

Exact task metadata

scale-ai/instance_protonmail__webclients-32ff10999a06455cb2147f6873d627456924ae13 · Task ID scale-ai/instance_protonmail__webclients-32ff10999a06455cb2147f6873d627456924ae13

No recorded stop reason

Reward1
Assertions
Steps
Tool calls
Model calls
Tokens in / out
Not passedFailed on this receipt

Scale ai/instance flipt io flipt a0cbc0cb65ae601270bdbe3f5313e2dfd49c80e4

Did not complete the recorded checks. The record does not establish a cause.

Exact task metadata

scale-ai/instance_flipt-io__flipt-a0cbc0cb65ae601270bdbe3f5313e2dfd49c80e4 · Task ID scale-ai/instance_flipt-io__flipt-a0cbc0cb65ae601270bdbe3f5313e2dfd49c80e4

No recorded stop reason

Reward0
Assertions
Steps
Tool calls
Model calls
Tokens in / out
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
20.0%
10 / 50
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly

Published reference

Results reported by model developers and benchmark authors.

10 rows
#Model / recipeReported scoreRuns onEvidence
01Claude Fable 5Published reference · Reported evaluation
80.0%
80.0%
Published reference↗ Published resultPublished by OpenAI
01Claude Fable 5Anthropic system card · Anthropic evaluation
80.0%
80.0%
Published reference↗ Published resultAnthropic
03Claude Opus 4.8Published reference · Reported evaluation
69.2%
69.2%
Published reference↗ Published resultPublished by OpenAI
03Claude Opus 4.8Anthropic system card · Adaptive thinking · max
69.2%
69.2%
Published reference↗ Published resultAnthropic
05GPT-5.6 SolPublished reference · Reported evaluation
64.6%
64.6%
Published reference↗ Published resultPublished by OpenAI
06GPT-5.6 TerraPublished reference · Reported evaluation
63.4%
63.4%
Published reference↗ Published resultPublished by OpenAI
07GPT-5.6 LunaPublished reference · Reported evaluation
62.7%
62.7%
Published reference↗ Published resultPublished by OpenAI
08Qwen 3.8 27BBF16 reference · Reported evaluation
61.7%
61.7%
Published reference↗ Published resultQwen
09Gemini 3.1 Pro PreviewPublished reference · Reported evaluation
54.2%
54.2%
Published reference↗ Published resultPublished by OpenAI
09Gemini 3.1 Pro PreviewGoogle evaluation · High thinking · Google scaffold
54.2%
54.2%
Published reference↗ Published resultGoogle DeepMind