LocalsOnlyevaluationsHow to read this

GLM-5.3-Flash

Setups

Unspecified references

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

ON THIS SETUP · SOURCE-REPORTED

Z.ai reference

Setup unspecified · Reported evaluation

2 benchmark results
Coding

Terminal-Bench

Multi-step tasks in a terminal environment.

Scoring & limits

Version 2.1 · sample size not reported

This is the source’s reported aggregate. Its metric, agent configuration, and evaluation protocol determine how it should be interpreted. Read the original grading definition before comparing completion rates across setups.

84.3%Sample size not reported
Published · 2026-08-26Inspect evidence ↗
Office & CRM

AutomationBench

Zapier’s public 600-task set. Strict, all-assertions completion at the highest reported effort.

Scoring & limits

Published public evaluation · 600-task set

A workflow is complete only when every assertion passes. Mean reward gives partial credit for progress; it is not the share of jobs finished. Task sets and evaluation budgets can differ between receipts.

Fully completed48.8%Sample size not reported
Published · 2026-08-26Inspect evidence ↗

Task traces are not available for this setup.

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
Reported evaluation
Model host
Not recorded
Weight artifact
Not recorded
Weight revision
Not recorded

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.