τ²-Bench Tool-Agent-User Evaluation

ModelsModels with a result.132
Top resultBest result on this test.99.1%GLM-5.2
Top-3 spreadPoints from first to third.0.6 pts
YearYear of release.2025

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

132 results
#ModelResultThe score from the source.
01GLM-5.299.1%
02GPT-5.498.9%
07GLM-598.2%
08GPT-5.598.0%
09GLM-5.197.7%
09Grok 4.397.7%
13GLM-4.795.9%
35Muse Spark91.5%
42MiniMax M388.9%
54GPT-5.284.8%
64GPT-5.181.9%
65o380.7%
71GLM-4.676.9%
73Grok 474.9%
74Qwen3 Max74.3%
74K-Exaone74.3%
82o162.6%
83Kimi K261.1%
90GPT-4.147.1%
100DeepSeek-R136.5%
101Gemma 4 12B36.3%
103Sarvam 30B34.5%
106o3-mini28.7%
108GPT-4o25.1%
111DeepSeek V322.8%
114Gemma 4 E4B20.8%
114Gemma 4 E2B20.8%
120GPT-4.1 nano17.3%
125Nova Pro14.0%
128Gemma 3 27B10.5%
131Phi-40.0%