DeepSearchQA

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

  • In index
  • Agentic
ModelsModels with a result.18
Top resultBest result on this test.96.0%Atria Dawn Preview
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.N/A

Result and price

020406080100$0.1$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

18 results
#ModelResultThe score from the source.
02Kimi K395.0%
12Muse Spark74.8%
15GPT-5.473.6%
17Grok 4.2062.8%