LocalsOnlyevaluationsHow to read this

GPT-6 Sol

Setups

Published full-model reference, then measured local recipes.

Recipe by benchmark scores. Published reference is not a rank. Blank cells have no cite.
RecipeAuto
OpenAI comparison table / publishedPublished reference33.2

The published full-model reference is the top row — or the dashed gold tick on Graphic — not rank #1. A blank cell, missing tick, or missing bead has no recorded cite. Scores are never invented.

ON THIS SETUP · SOURCE-REPORTED

OpenAI comparison table

Cloud / API · OpenAI · xhigh

1 benchmark result
Office & CRM

AutomationBench · public 600 ↗

Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.

Scoring & limits

Published public evaluation · 600-task set

A workflow is complete only when every assertion passes. Mean reward gives partial credit for progress; it is not the share of jobs finished. Task sets and evaluation budgets can differ between receipts.

Fully completed33.2%Sample size not reported
Published · 2026-09-22Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
OpenAI · xhigh
Model host
Provider managed · physical host not recorded
Weight artifact
Provider managed
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.