Instruction Following Benchmark

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

  • In index
  • Instruction
ModelsModels with a result.42
Top resultBest result on this test.85.0%MAI-Thinking-1
Top-3 spreadPoints from first to third.2.8 pts
YearYear of release.2025

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

42 results
#ModelResultThe score from the source.
06Grok 4.381.3%
10Inkling79.8%
19A.X K275.9%
36ZAYA1-8B52.6%