Published reference2026-06-09
SWE-bench Verified
swe-bench-verified · 500 tasks. Compare only matched task sets and protocols.
Claude Opus 4.8 · Anthropic system card · Anthropic · Anthropic evaluation
↗ Published resultSWE-bench Verified
Claude Fable 5 and Mythos 5 System Card, Table 8.1.A and section 8.2 ↗Tokens per second · see measurement basis
As recorded on this receipt
Run configuration & provenance +
- Context
- Not recorded
- Temperature
- Not recorded
- Wall time
- Not recorded
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- Not reported
- Dataset revision
- Not reported
- Harness host
- Not reported
- Model host
- Published reference
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
Anthropic standard configuration · adaptive thinking max effort · default sampling · 5-trial average. n=500. Thinking blocks included.
Table 8.1.A lists Claude Opus 4.8 SWE-bench Verified 88.6. Section 8.2: 500-problem subset, average over five trials, standard configuration, thinking blocks included. Anthropic primary cite — not an OpenAI comparison-table figure. Not a LocalsOnly desk run.
Run ID: research-claude-opus-4.8-claude-code-anthropic-system-card-swe-bench-verifiedTask explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →