Office & CRMOffice completion · public
AutomationBench*
* Published public evaluation · 600-task set. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
Can it actually finish the office workflow?
Published evaluation. No local run on this task set.
See published results ↓How scoring works ↓THE QUESTION IT ANSWERS
Can it actually finish the office workflow?
Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.
- 01Business workflow→
- 02Operate connected tools→
- 03Final-state assertions
Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
What this benchmark measures
Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.
Completing a job means satisfying all assertions against the final system state. A confident completion message or partial progress does not necessarily mean the workflow succeeded.
AutomationBench paper ↗How a task earns its score
This board has published cites only. There is no local task grid to inspect.
Each published row is a source-reported headline. Protocols can differ across cites.
A passed task and a failed one
Results by subject or work category
Results & published references
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Published reference
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ✺Claude Opus 5AutomationBench · public · Native API · maximum effort | 50.3% | Published reference | ↗ Published resultZapier | ↗ |
| 02 | ◎GLM-5.3-FlashPublished reference · Reported evaluation | 48.8% | Published reference | ↗ Published resultZ.ai | ↗ |
| 03 | ✺Claude Fable 5AutomationBench · public · Native API · maximum effort | 46.2% | Published reference | ↗ Published resultZapier | ↗ |
| 04 | ◎GPT-5.6 SolAutomationBench · public · Native API · maximum effort | 45.8% | Published reference | ↗ Published resultZapier | ↗ |
| 05 | ✺Claude Opus 4.8AutomationBench · public · Native API · maximum effort | 41.0% | Published reference | ↗ Published resultZapier | ↗ |
| 06 | ◎GPT-5.6 TerraAutomationBench · public · Native API · maximum effort | 37.2% | Published reference | ↗ Published resultZapier | ↗ |
| 07 | ◎DeepSeek V4 Flash 0731Published reference · Reported evaluation | 25.1% | Published reference | ↗ Published resultDeepSeek | ↗ |