These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.
This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.
Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.
Scoring & limits
Published public evaluation · 600-task set
A workflow is complete only when every assertion passes. Mean reward gives partial credit for progress; it is not the share of jobs finished. Task sets and evaluation budgets can differ between receipts.