Instruction-Following Eval

A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.

ModelsModels with a result.33
Top resultBest result on this test.95.0%Qwen3.5-27B
Top-3 spreadPoints from first to third.0.2 pts
YearYear of release.2023

Result and price

020406080100$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

33 results
#ModelResultThe score from the source.
07o3-mini93.9%
11GLM-592.6%
14o192.2%
20GPT-4.187.4%
23ZAYA1-8B85.6%