LongBench v2 authors
LongBench v2
A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.
ModelsModels with a result.14
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2025
Result and price
Results
14 results#ModelResultThe score from the source.