Build pmars
Completed the recorded checks.
Exact task metadata
build-pmars · Task ID build-pmars
No recorded stop reason
CodingTerminal
Version 2.1 · 89 tasks. Exact task sets and protocols remain separate.
Can it finish a task in a real terminal?
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB
Inspect the measured run ↗How scoring works ↓THE QUESTION IT ANSWERS
Measures whether an agent can navigate an environment, execute commands, and complete a technical task.
Interpret the score The harness, tools, runtime, and context budget all contribute to the outcome.
BEYOND THE HEADLINE
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.
Inspect individual task outcomes →Multi-step tasks in a terminal environment.
The paper separates execution, coherence and verification failures. A task can fail because the environment or tools break, because the agent loses track of the goal, or because it fails to check its work.
Terminal-Bench research ↗Each scored item is a pass or fail in a terminal environment.
The LocalsOnly measured run is Harbor-style: per-task pass and, when present, an issue (for example AgentTimeoutError). A setup issue is an execution failure on the receipt, not a separate qualitative score.
Published rows are vendor or bench cites. They do not share the measured task grid.
Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-02. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.
passed vs AgentTimeoutError
Completed the recorded checks.
build-pmars · Task ID build-pmars
No recorded stop reason
Stopped at the recorded time limit.
adaptive-rejection-sampler · Task ID adaptive-rejection-sampler
AgentTimeoutError · Agent execution timed out after 900.0 seconds
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Results from the exact local setup shown below.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BUD-Q4_K_XL · llama.cpp | 30.3% | GB10 · 128GB | ● Tested by LocalsOnlyLocalsOnly | ↗ |
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ◎GPT-5.6 SolSol Ultra · four agents · Four-agent evaluation | 91.9% | Published reference | ↗ Published resultOpenAI | ↗ |
| 02 | ◎GPT-5.6 SolPublished reference · Reported evaluation | 88.8% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 03 | ◎GPT-5.6 TerraPublished reference · Reported evaluation | 87.4% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 04 | ◎GPT-5.6 LunaPublished reference · Reported evaluation | 84.7% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 05 | ◎GLM-5.3-FlashPublished reference · Reported evaluation | 84.3% | Published reference | ↗ Published resultZ.ai | ↗ |
| 05 | ✺Claude Fable 5Anthropic system card · Anthropic evaluation | 84.3% | Published reference | ↗ Published resultAnthropic | ↗ |
| 07 | ◎GLM-5.3-Flash NVFP4Published reference · Reported evaluation | 83.2% | Published reference | ↗ Published resultNVIDIA | ↗ |
| 08 | ✺Claude Fable 5Published reference · Reported evaluation | 83.1% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 09 | ◎DeepSeek V4 Flash 0731Published reference · Reported evaluation | 82.7% | Published reference | ↗ Published resultDeepSeek | ↗ |
| 10 | ✺Claude Opus 4.8Published reference · Reported evaluation | 78.9% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 11 | ✺Claude Opus 4.8Anthropic system card · Adaptive thinking · max | 74.6% | Published reference | ↗ Published resultAnthropic | ↗ |
| 12 | ✳Qwen 3.8 27BBF16 reference · Reported evaluation | 73.0% | Published reference | ↗ Published resultQwen | ↗ |
| 13 | ✦Gemini 3.1 Pro PreviewPublished reference · Reported evaluation | 70.7% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |