LocalsOnlyevaluationsHow to read this
← Tests

CodingSoftware engineering

SWE-bench Verified

swe-bench-verified · 500 tasks. Exact task sets and protocols remain separate.

Can it fix an issue in an existing codebase?

Published evaluation. No local run on this task set.

See published results ↓How scoring works ↓

THE QUESTION IT ANSWERS

Can it fix an issue in an existing codebase?

SWE-bench Verified: 500 human-verified repository issues. This is the public-comparable software-engineering headline. Agent and harness belong on the cite.

  1. 01Repository + issue
  2. 02Inspect and change code
  3. 03Repository tests

Interpret the score LocalsOnly has not run Verified yet. Empty MEASURED is unknown, not zero. SWE-bench Pro n=50 receipts are archived footnotes, not this score.

What this test is

What this benchmark measures

500 human-verified repository issues. Public-comparable headline for repo software engineering. Agent and harness belong on the cite.

SWE-bench Verified is a 500-issue subset that human engineers marked as solvable. A resolved score is a patch that passes the repository tests. Agent, harness, and attempt policy belong with the cite. This is not SWE-bench Pro.

SWE-bench paper
How it grades

How a task earns its score

This board has published cites only. There is no local task grid to inspect.

Each published row is a source-reported headline. Protocols can differ across cites.

SuiteSWE-bench Verified
Published cite4 published rows
Published citeVendor or bench number; no local task grid
Follow one task · this measured recipe

A passed task and a failed one

This board has no LocalsOnly measured receipt, so there is no task example to inspect. Published rows are source cites only.
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Published reference

Results reported by model developers and benchmark authors.

4 rows
#Model / recipeReported scoreRuns onEvidence
01Claude Fable 5Anthropic system card · Anthropic evaluation
95.0%
95.0%
Published reference↗ Published resultAnthropic
02Claude Opus 4.8Anthropic system card · Anthropic evaluation
88.6%
88.6%
Published reference↗ Published resultAnthropic
03Gemini 3.1 Pro PreviewGoogle evaluation · High thinking · Google scaffold
80.6%
80.6%
Published reference↗ Published resultGoogle DeepMind
04DeepSeek V4 Flash 0731Published reference · Reported evaluation
79.0%
79.0%
Published reference↗ Published resultDeepSeek