CodingSoftware engineering
SWE-bench Verified
swe-bench-verified · 500 tasks. Exact task sets and protocols remain separate.
Can it fix an issue in an existing codebase?
Published evaluation. No local run on this task set.
See published results ↓How scoring works ↓THE QUESTION IT ANSWERS
Can it fix an issue in an existing codebase?
SWE-bench Verified: 500 human-verified repository issues. This is the public-comparable software-engineering headline. Agent and harness belong on the cite.
- 01Repository + issue→
- 02Inspect and change code→
- 03Repository tests
Interpret the score LocalsOnly has not run Verified yet. Empty MEASURED is unknown, not zero. SWE-bench Pro n=50 receipts are archived footnotes, not this score.
What this benchmark measures
500 human-verified repository issues. Public-comparable headline for repo software engineering. Agent and harness belong on the cite.
SWE-bench Verified is a 500-issue subset that human engineers marked as solvable. A resolved score is a patch that passes the repository tests. Agent, harness, and attempt policy belong with the cite. This is not SWE-bench Pro.
SWE-bench paper ↗How a task earns its score
This board has published cites only. There is no local task grid to inspect.
Each published row is a source-reported headline. Protocols can differ across cites.
A passed task and a failed one
Results by subject or work category
Results & published references
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Published reference
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ✺Claude Fable 5Anthropic system card · Anthropic evaluation | 95.0% | Published reference | ↗ Published resultAnthropic | ↗ |
| 02 | ✺Claude Opus 4.8Anthropic system card · Anthropic evaluation | 88.6% | Published reference | ↗ Published resultAnthropic | ↗ |
| 03 | ✦Gemini 3.1 Pro PreviewGoogle evaluation · High thinking · Google scaffold | 80.6% | Published reference | ↗ Published resultGoogle DeepMind | ↗ |
| 04 | ◎DeepSeek V4 Flash 0731Published reference · Reported evaluation | 79.0% | Published reference | ↗ Published resultDeepSeek | ↗ |