LocalsOnlyevaluationsHow to read this
← Model evidence

Published reference2026-09-22

AutomationBench · public 600*

Published public evaluation · 600-task set. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

Claude Opus 5.5 · Published reference · Published by Anthropic · Adaptive thinking · max effort

↗ Published result
Published by Anthropic

AutomationBench

Introducing Claude Opus 5.5
Reported score40.0%
Throughput

Tokens per second · see measurement basis

Concurrency

As recorded on this receipt

Published reference. This is a source-reported result, not a test run by this lab. Anthropic launch table: Opus 5.5 AutomationBench 40.0% at adaptive thinking max effort. Zapier ran the early-access evaluation without fallback models, so safeguard interventions counted as failures. Not the LocalsOnly local-300 sample. Terminal-Bench 4.0 and SWE-bench Pro on the card are not Terminal-Bench 2.1 or SWE-bench Verified, so those cells stay empty. https://www.anthropic.com/claude-opus-5-5
Run configuration & provenance +
Context
Not recorded
Temperature
Not recorded
Wall time
Not recorded
Weight file
Provider-managed / not reported
Revision
Not reported
Reasoning
Not reported
Agent version
Not reported
Output token limit
Not reported
Dataset revision
Not reported
Harness host
Not reported
Model host
Published reference
Input tokens
Not reported
Output tokens
Not reported
Cached input tokens
Not reported
Agent turn / step limit
Not reported
Concurrency
Not reported
Weight file size
Not reported

Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.

Claude Code · adaptive thinking · max effort · Claude Opus 5.5. Zapier early-access AutomationBench, no fallback. API id claude-opus-5-5. Not a LocalsOnly desk run.

Anthropic launch table: Opus 5.5 AutomationBench 40.0% at adaptive thinking max effort. Zapier ran the early-access evaluation without fallback models, so safeguard interventions counted as failures. Not the LocalsOnly local-300 sample. Terminal-Bench 4.0 and SWE-bench Pro on the card are not Terminal-Bench 2.1 or SWE-bench Verified, so those cells stay empty. https://www.anthropic.com/claude-opus-5-5

Run ID: research-claude-opus-5.5-claude-code-published-reference-automationbench-public
Tasks

Task explorer

Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.

0 task records

No task records in this receipt.

The headline or summary is available above. Individual outcomes have not been provided.

Explore test coverage →