Task ID 2848
Completed the recorded checks.
Exact task metadata
2848 · Task ID 2848
No recorded stop reason
CodingCode generation
Version 6.0 · 100 tasks. Exact task sets and protocols remain separate.
Can it turn a problem into working code?
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB
Inspect the measured run ↗How scoring works ↓THE QUESTION IT ANSWERS
Useful for understanding self-contained coding ability, where correctness can be checked by running the answer.
Interpret the score Version and problem window matter. This is different from maintaining an existing codebase.
BEYOND THE HEADLINE
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.
Inspect individual task outcomes →Solve programming problems with executable tests.
Time-indexed programming problems support generation, repair, execution and prediction tests. Version, time window and scenario matter when interpreting a result.
LiveCodeBench paper ↗The LocalsOnly measured receipt records each task as pass or fail.
A task can also carry a setup issue recorded on the receipt (for example AgentTimeoutError).
17 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.
Published rows are vendor or bench cites. They do not share the measured task grid.
Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-04. Both from leetcode or numeric. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.
passed vs AgentTimeoutError
Completed the recorded checks.
2848 · Task ID 2848
No recorded stop reason
Stopped at the recorded time limit.
2808 · Task ID 2808
AgentTimeoutError · Agent execution timed out after 360.0 seconds
AgentTimeoutError with reward 0.0
Ranked by this file’s recorded pass rate. A blank rate is not a zero. Names come from the receipt.
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Results from the exact local setup shown below.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BUD-Q4_K_XL · llama.cpp | 66.0% | GB10 · 128GB | ● Tested by LocalsOnlyLocalsOnly | ↗ |
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BBF16 reference · Reported evaluation | 90.3% | Published reference | ↗ Published resultQwen | ↗ |