ResearchClawBench

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

  • Not in index
  • Agentic
ModelsModels with a result.19
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

19 results
#ModelResultThe score from the source.
GLM-5.220.7%
MiniMax M319.8%
Qwen3.7 Max18.7%
GLM-5.118.2%
Kimi K2.618.0%
Qwen3.6 Plus18.0%
GPT-5.517.0%
MiMo-V2.516.9%
GPT-5.415.3%
MiMo-V2-Pro15.3%
Qwen3.5 397B14.2%
Kimi K2.514.0%
Grok 4.113.5%
Grok 4.312.4%