Task ID 0
Completed the recorded checks.
Exact task metadata
0 · Task ID 0
No recorded stop reason
ReasoningScientific reasoning
Diamond · version 1.0 · 198 tasks. Exact task sets and protocols remain separate.
Can it reason through difficult science?
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB
Inspect the measured run ↗How scoring works ↓THE QUESTION IT ANSWERS
Tests demanding scientific knowledge and reasoning in biology, chemistry, and physics.
Interpret the score A correct multiple-choice answer does not establish research or experimental competence.
BEYOND THE HEADLINE
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.
Inspect individual task outcomes →Graduate-level science questions.
Expert-written biology, chemistry and physics questions measure difficult scientific question answering. They do not establish practical experimental or research competence.
GPQA paper ↗The LocalsOnly measured receipt records each task as pass or fail.
A task can also carry a setup issue recorded on the receipt (for example AgentTimeoutError).
0 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.
Published rows are vendor or bench cites. They do not share the measured task grid.
Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-04. Both from gpqa-diamond. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.
passed vs AgentTimeoutError
Completed the recorded checks.
0 · Task ID 0
No recorded stop reason
Stopped at the recorded time limit.
56 · Task ID 56
AgentTimeoutError
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Results from the exact local setup shown below.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BUD-Q4_K_XL · llama.cpp | 73.2% | GB10 · 128GB | ● Tested by LocalsOnlyLocalsOnly | ↗ |
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ◎GPT-6 AstraPublished reference · Reported evaluation | 96.0% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 02 | ✦Gemini 3.8 FlashPublished reference · Reported evaluation | 95.3% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 03 | ◎GPT-5.6 SolPublished reference · Reported evaluation | 94.6% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 04 | ✦Gemini 3.1 Pro PreviewPublished reference · Reported evaluation | 94.3% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 04 | ✦Gemini 3.1 Pro PreviewGoogle evaluation · High thinking · Google scaffold | 94.3% | Published reference | ↗ Published resultGoogle DeepMind | ↗ |
| 06 | ✺Claude Fable 5.1Published reference · Reported evaluation | 93.7% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 06 | ✺Claude Opus 5Published reference · Reported evaluation | 93.7% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 08 | ✺Claude Opus 4.8Anthropic system card · Adaptive thinking · max | 93.6% | Published reference | ↗ Published resultAnthropic | ↗ |
| 09 | ◎GPT-5.6 TerraPublished reference · Reported evaluation | 92.9% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 10 | ✺Claude Fable 5Published reference · Reported evaluation | 92.6% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 11 | ◎GPT-5.6 LunaPublished reference · Reported evaluation | 92.3% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 12 | ◎GLM-5.3-Flash NVFP4Published reference · Reported evaluation | 92.1% | Published reference | ↗ Published resultNVIDIA | ↗ |
| 13 | ✺Claude Opus 4.8Published reference · Reported evaluation | 92.0% | Published reference | ↗ Published resultPublished by OpenAI | ↗ |
| 14 | ✳Qwen 3.8 27BBF16 reference · Reported evaluation | 89.2% | Published reference | ↗ Published resultQwen | ↗ |