LocalsOnly · measured2026-10-01
IFBench
Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.
GLM-5.3-Flash · EXL3 TR3 4bpw · 2× GB10 (DGX Spark head + ASUS GX10 worker) · 256GB total · vllm
Needs two desk boxes: DGX Spark (serve head) + ASUS GX10 (tensor-parallel worker), 128GB each.
● Tested by LocalsOnlyTokens per second · see measurement basis
As recorded on this receipt
- Weights are Mia-AiLab EXL3 TR3 4bpw (community quant) with a DFlash2 speculator (k=7), fp8 KV cache; not Z.ai's official weights.
- Server-default sampling (generation_config temp 1.0 / top_p 0.95) unless overridden by the request; this run's requests set temperature 0, top_p stays at the server default. Thinking on.
- Reply budget max_tokens 90,000 (DeepSeek V4 Flash used 131,072) because the serve is capped at max_model_len 100,000.
- 18 of 300 responses came back empty (finish_reason length: thinking used the full 90,000-token budget before any answer) and 1 request (key 191) timed out during a serve outage; the key 191 re-run was one of the 18 empties. All 18 score as failures: 282 responses carried an answer and were scored on content, 18 empties scored as fails.
- Run interrupted once by a DGX Spark / ASUS GX10 outage on 2026-09-30 and resumed from the checkpoint at 192 rows (about 3.5 h lost to the outage and restart).
- Single-turn IFBench_test, 300 prompts, no Harbor (generate.py then IFBench run_eval). Reasoning stripped before scoring. Headline is prompt-level loose accuracy, as on the Flash-Next IFBench receipt.
- Needs two desk boxes: serve head DGX Spark spark-5c2f (:8888) + tensor-parallel worker ASUS GX10 gx10-e789 (2× GB10, 256GB total). Not a single 128GB desk box.
- Not the published IFBench card protocol; Z.ai publishes no IFBench figure for GLM-5.3-Flash, so no vendor cite is compared here.
Run configuration & provenance +
- Context
- 100,000
- Temperature
- 0
- Wall time
- Not recorded
- Revision
- 25a44fdbf16862a46b7cc9921142c6c81350af2f
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- 90,000
- Dataset revision
- IFBench_test single-turn n=300
- Harness host
- mac
- Model host
- 2× GB10 (DGX Spark head + ASUS GX10 worker) · 256GB total
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- 175.72 GB
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
IFBench_test single-turn n=300 (no multi-turn, no Harbor). Generate on Kris Mac mini → vLLM OpenAI-compat [local endpoint] (GLM-5.3-Flash-EXL3): DGX Spark spark-5c2f head + ASUS GX10 gx10-e789 TP worker (2× GB10 · 256GB total). vLLM EXL3 4bpw, TP=2, fp8 KV cache, DFlash2 speculator (k=7), max_model_len 100000. temp=0 (request), thinking ON, reasoning stripped before eval, max_tokens 90000. Headline = prompt-level loose accuracy. Job ifbench-mac-dgx-glm53flash-20260929a. MEASURED local IFBench; NOT the published frontier IFBench card row.
Run ID: measured-2x-gb10-glm53flash-exl3-4bpw-ifbench-2026-10-01Task explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →