MRCRv2

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

  • Not in index
  • Reasoning
ModelsModels with a result.9
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Result and price

020406080100$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

9 results
#ModelResultThe score from the source.
Qwen3.8 Max92.9%
Qwen3.7 Plus91.7%
Qwen3.7 Max90.4%
Gemma 4 12B43.4%