LocalsOnly · measured2026-09-08
IFBench
Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB · llama.cpp
● Tested by LocalsOnlyReported score42.3%127 of 300 tasks passed
Throughput—
Tokens per second · see measurement basis
Concurrency—
As recorded on this receipt
Run configuration & provenance +
- Context
- Not recorded
- Temperature
- Not recorded
- Wall time
- Not recorded
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- Not reported
- Dataset revision
- Not reported
- Harness host
- Not reported
- Model host
- ASUS GX10 · 128GB
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
Benchmark-specific agent harness
Run ID: measured-spark-1x-unsloth-ud-q4_k_xl-ifbench-test-thin-2026-09-08Tasks
Task explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
↗
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →