LongBench v2

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

ModelsModels with a result.14
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2025

Result and price

020406080100$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

14 results
#ModelResultThe score from the source.
Qwen3.8 Max66.3%
Qwen3.5 397B63.2%
Qwen3.6 Plus62.0%
Kimi K2.561.0%
GLM-560.8%
Qwen3.5-27B60.6%
Agents-A160.2%