AI-Needle

A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.

  • Not in index
  • Reasoning
ModelsModels with a result.4
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Results

4 results
#ModelResultThe score from the source.
Qwen3.5 397B68.7%
Qwen3.6 Plus68.3%
GLM-563.3%