Bowen Ye et al.
Claw-Eval
A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.
ModelsModels with a result.39
Top-3 spreadPoints from first to third.4.3 pts
YearYear of release.2026
Result and price
Results
39 results#ModelResultThe score from the source.