Jason Wei et al.
Measuring Short-Form Factuality in Large Language Models
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
ModelsModels with a result.3
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2024
Results
3 results#ModelResultThe score from the source.