LocalsOnlyevaluationsHow to read this
← Model shelfNVIDIA / Z.ai

GLM-5.3-Flash NVFP4

Setups

Unspecified references

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

ON THIS SETUP · SOURCE-REPORTED

NVIDIA reference

Setup unspecified · Reported evaluation

3 benchmark results
Coding

Terminal-Bench

Multi-step tasks in a terminal environment.

Scoring & limits

Version 2.1 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

83.2%Sample size not reported
Published · 2026-09-09Inspect evidence ↗
Instructions

IFBench

Follow verifiable instruction constraints.

Scoring & limits

Source-reported protocol · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

60.5%Sample size not reported
Published · 2026-09-09Inspect evidence ↗
Reasoning

GPQA Diamond

Graduate-level science questions.

Scoring & limits

Diamond · version 1.0 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

92.1%Sample size not reported
Published · 2026-09-09Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Reported evaluation
Model host
Not recorded
Weight artifact
Not recorded
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.