LocalsOnlyevaluationsHow to read this
← Model evidence

LocalsOnly · measured2026-09-08

IFBench

Loose prompt grading · 300 tasks. Compare only matched task sets and protocols.

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB · llama.cpp

● Tested by LocalsOnly
Reported score42.3%127 of 300 tasks passed
Throughput

Tokens per second · see measurement basis

Concurrency

As recorded on this receipt

Run configuration & provenance +
Context
Not recorded
Temperature
Not recorded
Wall time
Not recorded
Weight file
Provider-managed / not reported
Revision
Not reported
Reasoning
Not reported
Agent version
Not reported
Output token limit
Not reported
Dataset revision
Not reported
Harness host
Not reported
Model host
ASUS GX10 · 128GB
Input tokens
Not reported
Output tokens
Not reported
Cached input tokens
Not reported
Agent turn / step limit
Not reported
Concurrency
Not reported
Weight file size
Not reported

Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.

Benchmark-specific agent harness

Run ID: measured-spark-1x-unsloth-ud-q4_k_xl-ifbench-test-thin-2026-09-08
Tasks

Task explorer

Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.

0 task records

No task records in this receipt.

The headline or summary is available above. Individual outcomes have not been provided.

Explore test coverage →