LocalsOnly · measured2026-09-09
IFBench
Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.
Qwen 3.8 Flash-Next · UD-IQ3_XXS · GB10 · 128GB · llama.cpp
● Tested by LocalsOnlyWeight pin not recorded on this receipt. No download URL.
Tokens per second · see measurement basis
As recorded on this receipt
Run configuration & provenance +
- Context
- Not recorded
- Temperature
- 0
- Wall time
- Not recorded
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- Not reported
- Dataset revision
- IFBench_test single-turn n=300
- Harness host
- Not reported
- Model host
- DGX Spark · 128GB
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
IFBench_test single-turn n=300 (no multi-turn). Generate on Kris Mac → DGX Spark spark-5c2f llama.cpp OpenAI-compat :8888 (qwen3.8-flash-next UD-IQ3_XXS). temp=0, thinking ON, strip <think> before eval. Headline = prompt-level loose accuracy. max_tokens raised as needed (8192→16384→…; probe still empty at 131072). Job ifbench-mac-dgx-flash-20260908a. RECEIPT — local Flash-Next IFBench; NOT the published frontier IFBench card row. Unlike 27B thin 0.423 thinking-off. HF pin unknown/null. Hardware evidence id dgx-spark = spark-5c2f (GB10 · 128GB; not Asus gx10-e789).
Run ID: measured-dgx-spark-unsloth-ud-iq3_xxs-ifbench-2026-09-09Task explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →