LocalsOnlyevaluationsHow to read this
← Tests

ReasoningScientific reasoning

GPQA Diamond

Diamond · version 1.0 · 198 tasks. Exact task sets and protocols remain separate.

Can it reason through difficult science?

73.2%Recorded score · 145 / 198 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Can it reason through difficult science?

Tests demanding scientific knowledge and reasoning in biology, chemistry, and physics.

  1. 01Expert-written question
  2. 02Choose an answer
  3. 03Correctness

Interpret the score A correct multiple-choice answer does not establish research or experimental competence.

BEYOND THE HEADLINE

What happened across 198 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed145 73.2%
No credit53 26.8%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
What this test is

What this benchmark measures

Graduate-level science questions.

Expert-written biology, chemistry and physics questions measure difficult scientific question answering. They do not establish practical experimental or research competence.

GPQA paper
How it grades

How a task earns its score

The LocalsOnly measured receipt records each task as pass or fail.

A task can also carry a setup issue recorded on the receipt (for example AgentTimeoutError).

0 infrastructure outcomes are recorded on the receipt. They are not invented extra grades.

Published rows are vendor or bench cites. They do not share the measured task grid.

SuiteGPQA Diamond
Tasks198 on the measured receipt
PassRecorded pass on the task
FailRecorded fail on the task
Setup issueAgentTimeoutError
Follow one task · this measured recipe

A passed task and a failed one

Same LocalsOnly receipt on GB10 · 128GB. Qwen 3.8 27B · UD-Q4_K_XL · 2026-09-04. Both from gpqa-diamond. Assertions, steps, and tool calls are on the task record. There is no agent transcript on this page.

passed vs AgentTimeoutError

Passedgpqa-diamond

Task ID 0

Completed the recorded checks.

Exact task metadata

0 · Task ID 0

No recorded stop reason

Reward1
Assertions
Steps
Tool calls
Model calls
Tokens in / out2,689 / 904
Not passedgpqa-diamond

Task ID 56

Stopped at the recorded time limit.

Exact task metadata

56 · Task ID 56

AgentTimeoutError

Reward0
Assertions
Steps
Tool calls
Model calls
Tokens in / out113,985 / 17,930
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
73.2%
145 / 198
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly

Published reference

Results reported by model developers and benchmark authors.

14 rows
#Model / recipeReported scoreRuns onEvidence
01GPT-6 AstraPublished reference · Reported evaluation
96.0%
96.0%
Published reference↗ Published resultPublished by OpenAI
02Gemini 3.8 FlashPublished reference · Reported evaluation
95.3%
95.3%
Published reference↗ Published resultPublished by OpenAI
03GPT-5.6 SolPublished reference · Reported evaluation
94.6%
94.6%
Published reference↗ Published resultPublished by OpenAI
04Gemini 3.1 Pro PreviewPublished reference · Reported evaluation
94.3%
94.3%
Published reference↗ Published resultPublished by OpenAI
04Gemini 3.1 Pro PreviewGoogle evaluation · High thinking · Google scaffold
94.3%
94.3%
Published reference↗ Published resultGoogle DeepMind
06Claude Fable 5.1Published reference · Reported evaluation
93.7%
93.7%
Published reference↗ Published resultPublished by OpenAI
06Claude Opus 5Published reference · Reported evaluation
93.7%
93.7%
Published reference↗ Published resultPublished by OpenAI
08Claude Opus 4.8Anthropic system card · Adaptive thinking · max
93.6%
93.6%
Published reference↗ Published resultAnthropic
09GPT-5.6 TerraPublished reference · Reported evaluation
92.9%
92.9%
Published reference↗ Published resultPublished by OpenAI
10Claude Fable 5Published reference · Reported evaluation
92.6%
92.6%
Published reference↗ Published resultPublished by OpenAI
11GPT-5.6 LunaPublished reference · Reported evaluation
92.3%
92.3%
Published reference↗ Published resultPublished by OpenAI
12GLM-5.3-Flash NVFP4Published reference · Reported evaluation
92.1%
92.1%
Published reference↗ Published resultNVIDIA
13Claude Opus 4.8Published reference · Reported evaluation
92.0%
92.0%
Published reference↗ Published resultPublished by OpenAI
14Qwen 3.8 27BBF16 reference · Reported evaluation
89.2%
89.2%
Published reference↗ Published resultQwen