Published reference2026-07-09
Terminal-Bench
Version 2.1 · sample size not reported. Compare only matched task sets and protocols.
GPT-5.6 Sol · Sol Ultra · four agents · OpenAI · Four-agent evaluation
↗ Published resultOpenAI
Protocol as described by the publisher
GPT-5.6 Sol Ultra evaluation ↗Reported score91.9%
Throughput—
Tokens per second · see measurement basis
Concurrency—
As recorded on this receipt
Published reference. This is a source-reported result, not a test run by this lab. Ultra uses four agents. This is a separate configuration from the primary reference.
Run configuration & provenance +
- Context
- Not recorded
- Temperature
- Not recorded
- Wall time
- Not recorded
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- Not reported
- Dataset revision
- Not reported
- Harness host
- Not reported
- Model host
- Published reference
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
Provider evaluation
Ultra uses four agents. This is a separate configuration from the primary reference.
Run ID: research-gpt-5.6-sol-codex-max-published-ultra-terminal-bench-2.1Tasks
Task explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
↗
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →