LocalsOnlyevaluationsHow to read this
← Tests

InstructionsFollowing instructions

IFBench

Loose prompt grading · 300 tasks. Exact task sets and protocols remain separate.

Will it follow the precise instructions?

42.3%Recorded score · 127 / 300 tasks

Qwen 3.8 27B · UD-Q4_K_XL · GB10 · 128GB

Inspect the measured run ↗How scoring works ↓

THE QUESTION IT ANSWERS

Will it follow the precise instructions?

Looks beyond a plausible answer to whether a response obeys specific, verifiable requirements.

  1. 01Prompt + constraints
  2. 02Produce a response
  3. 03Constraint checks

Interpret the score Our local headline uses loose prompt-level grading. Strict grading and published protocols differ.

SETUP PERFORMANCE

A capability profile, test by test

Each bar uses its benchmark’s scoring rule. Higher is better.

050100%
IFBenchTested by LocalsOnly
42.3%

Open any bar for the score, sample size, and evaluation settings. A missing measurement is never shown as zero.

What this test is

What this benchmark measures

Instruction-following constraints. Local receipt uses a different protocol from the vendor card.

The benchmark tests generalization to 58 unfamiliar precise constraints. Strict and loose criteria answer different questions and must be labeled separately.

IFBench paper
How it grades

How a task earns its score

Instruction-following is reported as an aggregate. The local receipt uses a different protocol from the vendor card and has no per-task traces on this board.

Published rows are vendor or bench cites. They do not share the measured task grid.

SuiteIFBench
ReceiptNo per-task records
Headline / aggregateNo per-task records on this receipt
Follow one task · this measured recipe

A passed task and a failed one

The LocalsOnly measured receipt for this test has no per-task records to inspect. The aggregate result is available below.
Subject & category breakdown · this recipe

Results by subject or work category

This receipt does not record a subject or work-category breakdown. A suite-wide label cannot support a finer subject comparison. Outcome groups such as “passed” and “timeout” are not subjects.
Ranking

Results & published references

Place is assigned only among rows that share a comparable score on this board. Evaluation settings can differ across sources.

Tested by LocalsOnly

Results from the exact local setup shown below.

1 row
#Model / recipeReported scoreRuns onEvidence
Qwen 3.8 27BUD-Q4_K_XL · llama.cpp
127 / 300
GB10 · 128GB● Tested by LocalsOnlyLocalsOnly

Published reference

Results reported by model developers and benchmark authors.

2 rows
#Model / recipeReported scoreRuns onEvidence
01Qwen 3.8 27BBF16 reference · Reported evaluation
79.5%
79.5%
Published reference↗ Published resultQwen
02GLM-5.3-Flash NVFP4Published reference · Reported evaluation
60.5%
60.5%
Published reference↗ Published resultNVIDIA