Jeffrey Zhou et al.
Instruction-Following Eval
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
ModelsModels with a result.33
Top-3 spreadPoints from first to third.0.2 pts
YearYear of release.2023
Result and price
Results
33 results#ModelResultThe score from the source.