InstructionsFollowing instructions
IFBench
Loose prompt grading · 300 tasks. Exact task sets and protocols remain separate.
Will it follow the precise instructions?
Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB
Inspect the measured run ↗How scoring works ↓THE QUESTION IT ANSWERS
Will it follow the precise instructions?
Looks beyond a plausible answer to whether a response obeys specific, verifiable requirements.
- 01Prompt + constraints→
- 02Produce a response→
- 03Constraint checks
Interpret the score Our local headline uses loose prompt-level grading. Strict grading and published protocols differ.
SETUP PERFORMANCE
A capability profile, test by test
Each bar uses its benchmark’s scoring rule. Higher is better.
Open any bar for the score, sample size, and evaluation settings. A missing measurement is never shown as zero.
What this benchmark measures
Instruction-following constraints. Local receipt uses a different protocol from the vendor card.
The benchmark tests generalization to 58 unfamiliar precise constraints. Strict and loose criteria answer different questions and must be labeled separately.
IFBench paper ↗How a task earns its score
Instruction-following is reported as an aggregate. The local receipt uses a different protocol from the vendor card and has no per-task traces on this board.
Published rows are vendor or bench cites. They do not share the measured task grid.
A passed task and a failed one
Results by subject or work category
Results & published references
Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.
Tested by LocalsOnly
Results from the exact local setup shown below.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| — | ✳Qwen 3.8 27BUD-Q4_K_XL · llama.cpp | — | GB10 · 128GB | ● Tested by LocalsOnlyLocalsOnly | ↗ |
Published reference
Results reported by model developers and benchmark authors.
| # | Model / recipe | Reported score ↓ | Runs on | Evidence | |
|---|---|---|---|---|---|
| 01 | ✳Qwen 3.8 27BBF16 reference · Reported evaluation | 79.5% | Published reference | ↗ Published resultQwen | ↗ |
| 02 | ◎GLM-5.3-Flash NVFP4Published reference · Reported evaluation | 60.5% | Published reference | ↗ Published resultNVIDIA | ↗ |