LocalsOnlyevaluationsHow to read this
← Model evidence

LocalsOnly · measured2026-10-01

IFBench

Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.

GLM-5.3-Flash · EXL3 TR3 4bpw · 2× GB10 (DGX Spark head + ASUS GX10 worker) · 256GB total · vllm

Needs two desk boxes: DGX Spark (serve head) + ASUS GX10 (tensor-parallel worker), 128GB each.

● Tested by LocalsOnly
Download weights · Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw @ 25a44fdb ↗Serve recipe · MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks@f4970207 ↗
Reported score67.7%203 of 300 tasks passed
Throughput—

Tokens per second · see measurement basis

Concurrency—

As recorded on this receipt

0 infrastructure outcomes are recorded. Failure counts may overlap with task outcomes.
Read before comparing
  • Weights are Mia-AiLab EXL3 TR3 4bpw (community quant) with a DFlash2 speculator (k=7), fp8 KV cache; not Z.ai's official weights.
  • Server-default sampling (generation_config temp 1.0 / top_p 0.95) unless overridden by the request; this run's requests set temperature 0, top_p stays at the server default. Thinking on.
  • Reply budget max_tokens 90,000 (DeepSeek V4 Flash used 131,072) because the serve is capped at max_model_len 100,000.
  • 18 of 300 responses came back empty (finish_reason length: thinking used the full 90,000-token budget before any answer) and 1 request (key 191) timed out during a serve outage; the key 191 re-run was one of the 18 empties. All 18 score as failures: 282 responses carried an answer and were scored on content, 18 empties scored as fails.
  • Run interrupted once by a DGX Spark / ASUS GX10 outage on 2026-09-30 and resumed from the checkpoint at 192 rows (about 3.5 h lost to the outage and restart).
  • Single-turn IFBench_test, 300 prompts, no Harbor (generate.py then IFBench run_eval). Reasoning stripped before scoring. Headline is prompt-level loose accuracy, as on the Flash-Next IFBench receipt.
  • Needs two desk boxes: serve head DGX Spark spark-5c2f (:8888) + tensor-parallel worker ASUS GX10 gx10-e789 (2× GB10, 256GB total). Not a single 128GB desk box.
  • Not the published IFBench card protocol; Z.ai publishes no IFBench figure for GLM-5.3-Flash, so no vendor cite is compared here.
Run configuration & provenance +
Context
100,000
Temperature
0
Wall time
Not recorded
Revision
25a44fdbf16862a46b7cc9921142c6c81350af2f
Reasoning
Not reported
Agent version
Not reported
Output token limit
90,000
Dataset revision
IFBench_test single-turn n=300
Harness host
mac
Model host
2× GB10 (DGX Spark head + ASUS GX10 worker) · 256GB total
Input tokens
Not reported
Output tokens
Not reported
Cached input tokens
Not reported
Agent turn / step limit
Not reported
Concurrency
Not reported
Weight file size
175.72 GB

Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.

IFBench_test single-turn n=300 (no multi-turn, no Harbor). Generate on Kris Mac mini → vLLM OpenAI-compat [local endpoint] (GLM-5.3-Flash-EXL3): DGX Spark spark-5c2f head + ASUS GX10 gx10-e789 TP worker (2× GB10 · 256GB total). vLLM EXL3 4bpw, TP=2, fp8 KV cache, DFlash2 speculator (k=7), max_model_len 100000. temp=0 (request), thinking ON, reasoning stripped before eval, max_tokens 90000. Headline = prompt-level loose accuracy. Job ifbench-mac-dgx-glm53flash-20260929a. MEASURED local IFBench; NOT the published frontier IFBench card row.

Run ID: measured-2x-gb10-glm53flash-exl3-4bpw-ifbench-2026-10-01
Tasks

Task explorer

Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.

0 task records
↗

No task records in this receipt.

The headline or summary is available above. Individual outcomes have not been provided.

Explore test coverage →