Office & CRMOffice completion · held-out
AutomationBench*
* Published private evaluation · sample size not reported. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
Can it actually finish the office workflow?
Published evaluation. No local run on this task set.
See published results ↓How scoring works ↓THE QUESTION IT ANSWERS
Can it actually finish the office workflow?
Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.
- 01Business workflow→
- 02Operate connected tools→
- 03Final-state assertions
Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
What this benchmark measures
Zapier’s private held-out evaluation. Separate task set from the public benchmark and local 300-task sample.
How a task earns its score
This board has published cites only. There is no local task grid to inspect.
Each published row is a source-reported headline. Protocols can differ across cites.
A passed task and a failed one
Results by subject or work category
Results & published references
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Published reference
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ◎GPT-6 AstraAutomationBench · private · Native API · Max | 41.4% | Published reference | ↗ Published resultZapier | ↗ |
| 02 | ✦Gemini 3.8 FlashAutomationBench · private · Native API · Medium | 29.7% | Published reference | ↗ Published resultZapier | ↗ |
| 03 | ◎GPT-5.6 SolAutomationBench · private · Native API · Max | 28.8% | Published reference | ↗ Published resultZapier | ↗ |