LocalsOnlyevaluationsHow to read this
← Tests

CodingCode generation

LiveCodeBench

Version 6.0 · 100 tasks. Exact task sets and protocols remain separate.

Can it turn a problem into working code?

66.0%Recorded score · 66 / 100 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Can it turn a problem into working code?

Useful for understanding self-contained coding ability, where correctness can be checked by running the answer.

  1. 01Programming problem
  2. 02Generate a solution
  3. 03Executable tests

Interpret the score Version and problem window matter. This is different from maintaining an existing codebase.

BEYOND THE HEADLINE

What happened across 100 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed66 66.0%
No credit17 17.0%
Other / unscored17 17.0%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
What this test is

What this benchmark measures

Solve programming problems with executable tests.

Time-indexed programming problems support generation, repair, execution and prediction tests. Version, time window and scenario matter when interpreting a result.

LiveCodeBench paper
How it grades

How a task earns its score

The LocalsOnly measured receipt records each task as pass or fail.

A task can also carry a setup issue recorded on the receipt (for example AgentTimeoutError).

17 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.

Published rows are vendor or bench cites. They do not share the measured task grid.

SuiteLiveCodeBench
Tasks100 on the measured receipt
PassRecorded pass on the task
FailRecorded fail on the task
Setup issueAgentTimeoutError
Follow one task · this measured recipe

A passed task and a failed one

Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-04. Both from leetcode or numeric. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.

passed vs AgentTimeoutError

Passedleetcode or numeric

Task ID 2848

Completed the recorded checks.

Exact task metadata

2848 · Task ID 2848

No recorded stop reason

Reward1
Assertions
Steps
Tool calls
Model calls
Tokens in / out31,073 / 1,979
Not passedleetcode or numeric

Task ID 2808

Stopped at the recorded time limit.

Exact task metadata

2808 · Task ID 2808

AgentTimeoutError · Agent execution timed out after 360.0 seconds

Reward0
Assertions
Steps
Tool calls
Model calls
Tokens in / out28,655 / 1,756

AgentTimeoutError with reward 0.0

Subject & category breakdown · this recipe

Results by subject or work category

Ranked by this file’s recorded pass rate. A blank rate is not a zero. Names come from the receipt.

  1. codeforces like100.0%1/1
  2. atcoder abc80.0%32/40
  3. leetcode or numeric78.6%33/42
  4. atcoder arc0/0
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
66.0%
66 / 100
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly

Published reference

Results reported by model developers and benchmark authors.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BBF16 reference · Reported evaluation
90.3%
90.3%
Published reference↗ Published resultQwen