LocalsOnlyevaluationsHow to read this

Claude Opus 4.8

Setups

Cloud / API

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

Weight pin not recorded on this receipt. No download URL.

Recipe JSON ↓

ON THIS SETUP · SOURCE-REPORTED

Anthropic reference

Cloud / API · Anthropic evaluation

1 benchmark result
Coding

SWE-bench Verified

500 human-verified repository issues. Public-comparable headline for repo software engineering. Agent and harness belong on the cite.

Scoring & limits

swe-bench-verified · 500 tasks

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

88.6%500 evaluated tasks
Published · 2026-06-09Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Anthropic evaluation
Model host
Provider managed · physical host not recorded
Weight artifact
Provider managed
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.