LocalsOnlyevaluationsHow to read this
← Model shelfGoogle DeepMind

Gemini 3.8 Flash

Setups

Cloud / API

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

ON THIS SETUP · SOURCE-REPORTED

OpenAI comparison table

Cloud / API · Reported evaluation

1 benchmark result
Reasoning

GPQA Diamond

Graduate-level science questions.

Scoring & limits

Diamond · version 1.0 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

95.3%Sample size not reported
Published · 2026-09-03Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Reported evaluation
Model host
Provider managed · physical host not recorded
Weight artifact
Provider managed
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.