LocalsOnlyevaluationsHow to read this

Claude Opus 4.8

Setups

Cloud / API

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

ON THIS SETUP · SOURCE-REPORTED

Anthropic reference

Cloud / API · Adaptive thinking · max

3 benchmark results
Coding

SWE-bench Pro

Resolve issues in real software repositories.

Scoring & limits

Pro evaluation · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

69.2%Sample size not reported
Published · 2026-05-28Inspect evidence ↗
Coding

Terminal-Bench

Multi-step tasks in a terminal environment.

Scoring & limits

Version 2.1 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

74.6%Sample size not reported
Published · 2026-05-28Inspect evidence ↗
Reasoning

GPQA Diamond

Graduate-level science questions.

Scoring & limits

Diamond · version 1.0 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

93.6%Sample size not reported
Published · 2026-05-28Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Adaptive thinking · max
Model host
Provider managed · physical host not recorded
Weight artifact
Provider managed
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.