LocalsOnlyevaluationsHow to read this

Claude Opus 5.5

Setups

Published full-model reference, then measured local recipes.

Recipe by benchmark scores. Published reference is not a rank. Blank cells have no cite.
RecipeAuto
Anthropic comparison table / publishedPublished reference40.0

The published full-model reference is the top row — or the dashed gold tick on Graphic — not rank #1. A blank cell, missing tick, or missing bead has no recorded cite. Scores are never invented.

ON THIS SETUP · SOURCE-REPORTED

Anthropic comparison table

Cloud / API · Adaptive thinking · max effort

1 benchmark result
Office & CRM

AutomationBench · public 600

Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.

Scoring & limits

Published public evaluation · 600-task set

A workflow is complete only when every assertion passes. Mean reward gives partial credit for progress; it is not the share of jobs finished. Task sets and evaluation budgets can differ between receipts.

Fully completed40.0%Sample size not reported
Published · 2026-09-22Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Adaptive thinking · max effort
Model host
Provider managed · physical host not recorded
Weight artifact
Provider managed
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.