LocalsOnlyevaluationsHow to read this
← Tests

Office & CRMOffice completion · held-out

AutomationBench*

* Published private evaluation · sample size not reported. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

Can it actually finish the office workflow?

Published evaluation. No local run on this task set.

See published results ↓How scoring works ↓

THE QUESTION IT ANSWERS

Can it actually finish the office workflow?

Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.

  1. 01Business workflow
  2. 02Operate connected tools
  3. 03Final-state assertions

Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

What this test is

What this benchmark measures

Zapier’s private held-out evaluation. Separate task set from the public benchmark and local 300-task sample.

How it grades

How a task earns its score

This board has published cites only. There is no local task grid to inspect.

Each published row is a source-reported headline. Protocols can differ across cites.

SuiteAutomationBench
Published cite3 published rows
Published citeVendor or bench number; no local task grid
Follow one task · this measured recipe

A passed task and a failed one

This board has no LocalsOnly measured receipt, so there is no task example to inspect. Published rows are source cites only.
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Published reference

Results reported by model developers and benchmark authors.

3 rows
#Model / recipeReported scoreRuns onEvidence
01GPT-6 AstraAutomationBench · private · Native API · Max
41.4%
41.4%
Published reference↗ Published resultZapier
02Gemini 3.8 FlashAutomationBench · private · Native API · Medium
29.7%
29.7%
Published reference↗ Published resultZapier
03GPT-5.6 SolAutomationBench · private · Native API · Max
28.8%
28.8%
Published reference↗ Published resultZapier