Update contact phone · sales
Completed the recorded checks.
Exact task metadata
sales.update_contact_phone · Task ID 1
stop
Office & CRMOffice automation
* Local evaluation · 300 tasks. Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
Can it actually finish the office workflow?
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB
Inspect the measured run ↗How scoring works ↓THE QUESTION IT ANSWERS
Tests sales, marketing, operations, support, finance, and HR workflows. Completion requires every assertion to pass.
Interpret the score Task sets and protocols differ. A local sample is not established as a subset of the published public set; scores remain separate.
BEYOND THE HEADLINE
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.
Inspect individual task outcomes →Sales, marketing, operations, support, finance and HR workflows.
Fully completed is nPass / nTasks. Mean reward is partial credit and stays a separate labeled note.
Some local tasks also record a setup issue (for example aborted_rollout).
Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-05. Both from sales. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.
passed in 13 steps / 24 tools vs aborted_rollout
Completed the recorded checks.
sales.update_contact_phone · Task ID 1
stop
The attempt stopped before completion. The record does not establish why.
sales.create_important_draft · Task ID 9
length · aborted_rollout
in export summary.aborted_tasks; log abort line does not name the task, so exceed_context vs tool-JSON is not assigned per task
Ranked by this file’s recorded fully completed rate. A blank rate is not a zero. Names come from the receipt.
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Results from the exact local setup shown below.
| # | Model / recipe | Fully completed ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BUD-Q4_K_XL · llama.cpp | 18.3% | GB10 · 128GB | ● Tested by LocalsOnlyLocalsOnly | ↗ |