LocalsOnly · measured2026-09-12
IFBench
Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.
DeepSeek V4 Flash · UD-IQ3_XXS · GB10 · 128GB · llama.cpp
● Tested by LocalsOnlyWeight pin not recorded on this receipt. No download URL.
Recorded effort for this run. Wall time and token counts are not a dollar cost; hardware power and utilization also matter.
Tokens per second · see measurement basis
As recorded on this receipt
- Unsloth DeepSeek-V4-Flash UD-IQ3_XXS GGUF via llama.cpp (llama-server) on DGX Spark spark-5c2f :8888 (1x GB10, 128GB). Immutable HF repo/revision/filename were not recorded in the job artifacts, so the weight pin is left null.
- The reply budget was max_tokens 131072 (set from the start, never lowered), but the live serve n_ctx was 32768 (job launch note, 1x DGX Spark, ~15Gi MemAvailable after load). Prompt + reply could never exceed 32768, so the 131072 budget was not reachable in practice.
- 40 of 300 responses came back empty: finish_reason length with 32,620-32,749 completion tokens each (the n_ctx wall), all thinking, no answer after the reasoning strip. All 40 score as failures in the loose count (196/300). 260 responses carried an answer; 196 of those 260 passed loose (75.4% of answered, shown for context only, not the headline).
- Sampling: the request set temperature 0 with thinking on (chat_template_kwargs enable_thinking true). top_p / top_k / other llama-server defaults were not recorded for this run.
- Single uninterrupted generation pass (no resume, no retries: attempts=1 on all 300 rows, no HTTP errors), 2026-09-10 17:15 ET to 2026-09-12 00:22 ET. Reasoning stripped before scoring (all 300 replies carried a reasoning block). The 'max_tokens probe still empty at 131072' note in the job dir reflects the same n_ctx wall.
- Single-turn IFBench_test, 300 prompts, no Harbor (generate.py then IFBench run_eval). Headline is prompt-level loose accuracy, as on the Flash-Next and GLM-5.3-Flash IFBench receipts.
- Not the published IFBench card protocol; this is a LocalsOnly desk recipe (UD-IQ3_XXS quant), not DeepSeek's vendor figure. Do not compare to a vendor IFBench number.
- Recount from run_eval artifacts (responses-eval_results_loose/strict.jsonl): loose 196/300 = 0.6533, strict 186/300 = 0.6200, matching the job's own run_eval log.
Run configuration & provenance +
- Context
- 32,768
- Temperature
- 0
- Wall time
- 1867.0 minutes
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- 131,072
- Dataset revision
- IFBench_test single-turn n=300
- Harness host
- mac
- Model host
- DGX Spark · 128GB
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: sum(completion_tokens) / sum(per-request elapsed_s) over the 300 checkpoint rows (2,048,412 tokens / 112,018 s); includes the 40 empty length-truncated requests.
IFBench_test single-turn n=300 (no multi-turn, no Harbor). Generate on Kris Mac mini → DGX Spark spark-5c2f llama.cpp OpenAI-compat [local endpoint] (deepseek-v4-flash UD-IQ3_XXS). temp=0 (request), thinking ON, reasoning stripped before eval, max_tokens 131072 against a live n_ctx of 32768. Headline = prompt-level loose accuracy. Job ifbench-mac-dgx-dsv4flash-20260910a. MEASURED local IFBench; NOT the published frontier IFBench card row. NOT Asus gx10-e789.
Run ID: measured-dgx-spark-unsloth-ud-iq3_xxs-ifbench-dsv4flash-2026-09-12Task explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →