LocalsOnlyevaluationsHow to read this
← Tests

CodingTerminal

Terminal-Bench

Version 2.1 · 89 tasks. Exact task sets and protocols remain separate.

Can it finish a task in a real terminal?

30.3%Recorded score · 27 / 89 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Can it finish a task in a real terminal?

Measures whether an agent can navigate an environment, execute commands, and complete a technical task.

  1. 01Task + environment
  2. 02Use tools over many steps
  3. 03Task verification

Interpret the score The harness, tools, runtime, and context budget all contribute to the outcome.

BEYOND THE HEADLINE

What happened across 89 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed27 30.3%
No credit60 67.4%
Other / unscored2 2.2%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
What this test is

What this benchmark measures

Multi-step tasks in a terminal environment.

The paper separates execution, coherence and verification failures. A task can fail because the environment or tools break, because the agent loses track of the goal, or because it fails to check its work.

Terminal-Bench research
How it grades

How a task earns its score

Each scored item is a pass or fail in a terminal environment.

The LocalsOnly measured run is Harbor-style: per-task pass and, when present, an issue (for example AgentTimeoutError). A setup issue is an execution failure on the receipt, not a separate qualitative score.

Published rows are vendor or bench cites. They do not share the measured task grid.

SuiteTerminal-Bench
Tasks89 on the measured receipt
PassRecorded pass on the task
FailRecorded fail on the task
Setup issueAgentTimeoutError
Follow one task · this measured recipe

A passed task and a failed one

Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-02. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.

passed vs AgentTimeoutError

PassedPassed on this receipt

Build pmars

Completed the recorded checks.

Exact task metadata

build-pmars · Task ID build-pmars

No recorded stop reason

Reward1
Assertions
Steps
Tool calls
Model calls
Tokens in / out
Not passedFailed on this receipt

Adaptive rejection sampler

Stopped at the recorded time limit.

Exact task metadata

adaptive-rejection-sampler · Task ID adaptive-rejection-sampler

AgentTimeoutError · Agent execution timed out after 900.0 seconds

Reward0
Assertions
Steps
Tool calls
Model calls
Tokens in / out
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
30.3%
27 / 89
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly

Published reference

Results reported by model developers and benchmark authors.

13 rows
#Model / recipeReported scoreRuns onEvidence
01GPT-5.6 SolSol Ultra · four agents · Four-agent evaluation
91.9%
91.9%
Published reference↗ Published resultOpenAI
02GPT-5.6 SolPublished reference · Reported evaluation
88.8%
88.8%
Published reference↗ Published resultPublished by OpenAI
03GPT-5.6 TerraPublished reference · Reported evaluation
87.4%
87.4%
Published reference↗ Published resultPublished by OpenAI
04GPT-5.6 LunaPublished reference · Reported evaluation
84.7%
84.7%
Published reference↗ Published resultPublished by OpenAI
05GLM-5.3-FlashPublished reference · Reported evaluation
84.3%
84.3%
Published reference↗ Published resultZ.ai
05Claude Fable 5Anthropic system card · Anthropic evaluation
84.3%
84.3%
Published reference↗ Published resultAnthropic
07GLM-5.3-Flash NVFP4Published reference · Reported evaluation
83.2%
83.2%
Published reference↗ Published resultNVIDIA
08Claude Fable 5Published reference · Reported evaluation
83.1%
83.1%
Published reference↗ Published resultPublished by OpenAI
09DeepSeek V4 Flash 0731Published reference · Reported evaluation
82.7%
82.7%
Published reference↗ Published resultDeepSeek
10Claude Opus 4.8Published reference · Reported evaluation
78.9%
78.9%
Published reference↗ Published resultPublished by OpenAI
11Claude Opus 4.8Anthropic system card · Adaptive thinking · max
74.6%
74.6%
Published reference↗ Published resultAnthropic
12Qwen 3.8 27BBF16 reference · Reported evaluation
73.0%
73.0%
Published reference↗ Published resultQwen
13Gemini 3.1 Pro PreviewPublished reference · Reported evaluation
70.7%
70.7%
Published reference↗ Published resultPublished by OpenAI