LocalsOnlyevaluationsHow to read this
← Tests

Office & CRMOffice automation

AutomationBench*

* Local evaluation · 300 tasks. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

Can it actually finish the office workflow?

18.3%Fully completed · 55 / 300 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Can it actually finish the office workflow?

Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.

  1. 01Business workflow
  2. 02Operate connected tools
  3. 03Final-state assertions

Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.

BEYOND THE HEADLINE

What happened across 300 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed55 18.3%
Partial progress149 49.7%
No credit96 32.0%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
What this test is

What this benchmark measures

Sales, marketing, operations, support, finance and HR workflows.

How it grades

How a task earns its score

Fully completed is nPass / nTasks. Mean reward is partial credit and stays a separate labeled note.

Some local tasks also record a setup issue (for example aborted_rollout).

SuiteAutomationBench
Tasks300 on the measured receipt
Fully completednPass / nTasks on the receipt
Mean rewardPartial credit; labeled separately
Setup issueaborted_rollout
Follow one task · this measured recipe

A passed task and a failed one

Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-05. Both from sales. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.

passed in 13 steps / 24 tools vs aborted_rollout

Passedsales

Update contact phone · sales

Completed the recorded checks.

Exact task metadata

sales.update_contact_phone · Task ID 1

stop

Reward1
Assertions4 / 4
Steps13
Tool calls24
Model calls13
Tokens in / out79,340 / 3,160
Abortedsales

Create important draft · sales

The attempt stopped before completion. The record does not establish why.

Exact task metadata

sales.create_important_draft · Task ID 9

length · aborted_rollout

Reward0
Assertions0 / 2
Steps26
Tool calls98
Model calls26
Tokens in / out420,749 / 10,281

in export summary.aborted_tasks; log abort line does not name the task, so exceed_context vs tool-JSON is not assigned per task

Subject & category breakdown · this recipe

Results by subject or work category

Ranked by this file’s recorded fully completed rate. A blank rate is not a zero. Names come from the receipt.

  1. operations26.0%13/50 completed · reward 55.5%
  2. hr18.0%9/50 completed · reward 34.5%
  3. support18.0%9/50 completed · reward 62.8%
  4. finance16.0%8/50 completed · reward 37.8%
  5. marketing16.0%8/50 completed · reward 60.5%
  6. sales16.0%8/50 completed · reward 41.8%
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeFully completedRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
18.3%
55 / 300Fully completedMean reward 48.8%
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly