Inference
Explore
Leaderboard
Benchmarks
Methodology
24 Sept 2026
Benchmarks
/
τ²-Bench Tool-Agent-User Evaluation
Victor Barres et al.
τ²-Bench Tool-Agent-User Evaluation
In index
Agentic
Source ↗
Models
?
Models with a result.
132
Top result
?
Best result on this test.
99.1
%
GLM-5.2
Top-3 spread
?
Points from first to third.
0.6
pts
Year
?
Year of release.
2025
Result and price
0
20
40
60
80
100
$0.1
$0.3
$1
$3
$10
$30
$100
RESULT
PARETO
OUTPUT PRICE PER 1M TOKENS · LOG SCALE
Results
132 results
#
Model
Result
?
The score from the source.
01
GLM-5.2
99.1
%
02
GPT-5.4
98.9
%
03
Claude Fable 5
98.5
%
03
ST
Step 3.7 Flash
98.5
%
03
GLM-5-Turbo
98.5
%
03
GLM-5V-Turbo
98.5
%
07
GLM-5
98.2
%
08
GPT-5.5
98.0
%
09
GLM-5.1
97.7
%
09
Qwen3.6 Plus
97.7
%
09
Grok 4.3
97.7
%
12
DeepSeek V4 Pro 0813
96.2
%
13
MA
Kimi K2.6
95.9
%
13
Qwen 3.6 Max (preview)
95.9
%
13
MA
Kimi K2.5 (Reasoning)
95.9
%
13
MA
Kimi K2.5
95.9
%
13
GLM-4.7
95.9
%
18
Gemini 3.1 Pro
95.6
%
19
Gemini 3.5 Flash
95.3
%
19
Qwen3.6-35B-A3B
95.3
%
21
MiMo-V2-Pro
95.0
%
22
Qwen3.7 Max
94.7
%
23
Claude Opus 4.8
94.4
%
24
MiMo-V2.5-Pro
94.2
%
24
Qwen3.6-27B
94.2
%
24
Mistral Medium 3.5 128B
94.2
%
27
Qwen3.5-27B
93.9
%
28
Qwen3.5-122B-A10B
93.6
%
29
GPT-5.4 mini
93.4
%
30
Grok 4.1 Fast (Reasoning)
93.3
%
31
Qwen3.7 Plus
93.0
%
32
GPT-5.4 nano
92.5
%
33
Claude Opus 4.6 (Adaptive)
92.1
%
33
GPT-5.2-Codex
92.1
%
35
Muse Spark
91.5
%
36
MiMo-V2-Omni
91.2
%
37
MA
Kimi K2.7 Code
90.1
%
37
AA
Trinity-Large-Thinking
90.1
%
37
AA
Trinity-Large-Preview
90.1
%
40
Claude Opus 4.5 Thinking
89.5
%
41
Qwen3.5-35B-A3B
89.2
%
42
MiniMax M3
88.9
%
43
Claude Opus 4.7 (Adaptive)
88.6
%
44
LI
LFM2.5-8B-A1B
88.1
%
45
Gemini 3 Pro
87.1
%
46
GPT-5 (medium)
86.5
%
47
GPT-5.6 Terra
86.3
%
47
Claude Opus 4.5
86.3
%
47
UP
Solar Pro 3
86.3
%
50
GPT-5.3 Codex
86.0
%
50
IN
Ling 2.6 Flash
86.0
%
52
GPT-5.6 Sol
85.1
%
53
CO
Command A+
85.0
%
54
GPT-5.2
84.8
%
54
Claude Opus 4.6
84.8
%
54
GPT-5 (high)
84.8
%
54
MiniMax M2.7
84.8
%
58
Qwen3.5 397B
83.9
%
58
Qwen3.5 397B (Reasoning)
83.9
%
58
MiMo-V2-Flash
83.9
%
61
Nemotron 3 Ultra
83.3
%
62
GPT-5.1-Codex-Max
83.0
%
62
GPT-5.1-Codex
83.0
%
64
GPT-5.1
81.9
%
65
o3
80.7
%
66
IN
LLaDA2.2-flash
80.3
%
67
PM
Ternary Bonsai 2 27B
80.2
%
68
Claude Sonnet 4.6
79.5
%
69
DeepSeek V3.2
78.9
%
70
IN
Agents-A1-4B
78.2
%
71
GLM-4.6
76.9
%
72
Grok Code Fast 1
75.7
%
73
Grok 4
74.9
%
74
Qwen3 Max
74.3
%
74
LA
K-Exaone
74.3
%
76
Claude Opus 4.7
74.0
%
77
Claude 4.1 Opus Thinking
71.4
%
78
Nemotron 3 Super 100B
67.8
%
79
Grok 4 Fast (Reasoning)
65.8
%
79
GPT-OSS 120B
65.8
%
81
Grok 4.1 Fast
63.7
%
82
o1
62.6
%
83
MA
Kimi K2
61.1
%
84
GPT-OSS 20B
60.2
%
85
Gemma 4 31B
59.9
%
86
IN
LLaDA2.2-mini
57.5
%
87
Gemini 2.5 Pro
54.1
%
88
GPT-4.1 mini
52.9
%
89
Claude 4 Sonnet
52.3
%
90
GPT-4.1
47.1
%
91
SA
Sarvam 105B
46.8
%
92
GLM-4.5-Air
46.5
%
93
Nemotron 3 Nano Omni 30B A3B
45.3
%
94
Gemma 4 26B A4B
43.6
%
95
Gemini 3 Flash
43.3
%
96
Mistral Small 4
41.2
%
—
Mistral Small 4 (Reasoning)
41.2
%
97
Nemotron 3 Nano 30B
40.9
%
98
CO
North Mini Code
37.4
%
98
DeepSeek V3.1 (Reasoning)
37.4
%
100
DeepSeek-R1
36.5
%
101
Gemma 4 12B
36.3
%
102
DeepSeek V3.1
34.8
%
103
SA
Sarvam 30B
34.5
%
104
UP
Solar Pro 2
31.9
%
105
Mistral Large 2
30.7
%
106
o3-mini
28.7
%
107
FA
Ultravox v0.6 Llama 3.3 70B
26.6
%
108
GPT-4o
25.1
%
109
Mistral Large 3
24.6
%
110
Mistral Medium 3
24.3
%
111
DeepSeek V3
22.8
%
112
Qwen3-Omni-30B-A3B-Thinking
21.3
%
113
Claude 3 Haiku
21.1
%
114
Gemma 4 E4B
20.8
%
114
Gemma 4 E2B
20.8
%
116
LA
Exaone 4.0 1.2B
20.5
%
117
IB
Granite-4.0-H-1B
19.6
%
118
Llama 3.1 405B
19.0
%
119
Llama 4 Maverick
17.8
%
120
GPT-4.1 nano
17.3
%
121
Qwen3-Omni-30B-A3B-Instruct
16.4
%
122
Llama 4 Scout
15.5
%
123
Gemini 2.5 Flash
14.9
%
124
IB
Granite-4.0-H-350M
14.6
%
125
AM
Nova Pro
14.0
%
126
IB
Granite-4.0-350M
13.2
%
127
Nemotron Ultra 253B
11.4
%
128
Gemma 3 27B
10.5
%
129
LI
LFM2.5-VL-1.6B-Extract
8.5
%
130
LA
Exaone 4.0 32B
4.1
%
131
Phi-4
0.0
%
τ²-Bench Tool-Agent-User Evaluation results — Inference360