LocalsOnlyevaluationsHow to read this
← Tests

Office & CRMOffice completion · public

AutomationBench*

* Published public evaluation · 600-task set. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

Can it actually finish the office workflow?

Published evaluation. No local run on this task set.

See published results ↓How scoring works ↓

THE QUESTION IT ANSWERS

Can it actually finish the office workflow?

Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.

  1. 01Business workflow
  2. 02Operate connected tools
  3. 03Final-state assertions

Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

What this test is

What this benchmark measures

Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.

Completing a job means satisfying all assertions against the final system state. A confident completion message or partial progress does not necessarily mean the workflow succeeded.

AutomationBench paper
How it grades

How a task earns its score

This board has published cites only. There is no local task grid to inspect.

Each published row is a source-reported headline. Protocols can differ across cites.

SuiteAutomationBench
Published cite7 published rows
Published citeVendor or bench number; no local task grid
Follow one task · this measured recipe

A passed task and a failed one

This board has no LocalsOnly measured receipt, so there is no task example to inspect. Published rows are source cites only.
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Published reference

Results reported by model developers and benchmark authors.

7 rows
#Model / recipeReported scoreRuns onEvidence
01Claude Opus 5AutomationBench · public · Native API · maximum effort
50.3%
50.3%
Published reference↗ Published resultZapier
02GLM-5.3-FlashPublished reference · Reported evaluation
48.8%
48.8%
Published reference↗ Published resultZ.ai
03Claude Fable 5AutomationBench · public · Native API · maximum effort
46.2%
46.2%
Published reference↗ Published resultZapier
04GPT-5.6 SolAutomationBench · public · Native API · maximum effort
45.8%
45.8%
Published reference↗ Published resultZapier
05Claude Opus 4.8AutomationBench · public · Native API · maximum effort
41.0%
41.0%
Published reference↗ Published resultZapier
06GPT-5.6 TerraAutomationBench · public · Native API · maximum effort
37.2%
37.2%
Published reference↗ Published resultZapier
07DeepSeek V4 Flash 0731Published reference · Reported evaluation
25.1%
25.1%
Published reference↗ Published resultDeepSeek