Claw-Eval

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

ModelsModels with a result.39
Top resultBest result on this test.81.4%Ornith-1.5-397B
Top-3 spreadPoints from first to third.4.3 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

39 results
#ModelResultThe score from the source.
04MiniMax M374.5%
16Muse Spark63.8%
21GLM-5.162.3%
21MiMo-V2.562.3%
24GPT-5.460.3%
29GLM-557.7%