BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

  • In index
  • Agentic
ModelsModels with a result.44
Top resultBest result on this test.92.5%Atria Dawn Preview
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.N/A

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100$300RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

44 results
#ModelResultThe score from the source.
04Kimi K391.2%
13GPT-5.584.4%
16MiniMax M383.5%
21GPT-5.482.7%
24Inkling77.1%
29GLM-5.168.0%
32GPT-5.265.8%
40GLM-4.752.0%