GPQA Diamond ↗
Graduate-level science questions.
Scoring & limits
Diamond · version 1.0 · 198 tasks
Credit comes from answering the science question correctly. Multiple-choice accuracy does not establish research or experimental competence.
Rating: GPQA Diamond
These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.
ON THIS SETUP · MEASURED BY LOCALSONLY
GB10 · 128GB · llama.cpp
Graduate-level science questions.
Diamond · version 1.0 · 198 tasks
Credit comes from answering the science question correctly. Multiple-choice accuracy does not establish research or experimental competence.
BEYOND THE HEADLINE
Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.
Tested by LocalsOnly · UD-IQ3_XXS · GB10 · 128GB. Partial progress still requires further work before a task is complete.
Inspect individual task outcomes →Completed the recorded checks.
Assertion counts not recorded
Follow this exact task ↗0 · ID 0
gpqa-diamond · passed
Stopped at the recorded time limit.
Assertion counts not recorded
Follow this exact task ↗10 · ID 10
gpqa-diamond · AgentTimeoutError
Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.
Download this setup’s recipe JSON ↓