Instruction Following Benchmark
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
ModelsModels with a result.42
Top-3 spreadPoints from first to third.2.8 pts
YearYear of release.2025
Result and price
Results
42 results#ModelResultThe score from the source.