Terminal-Bench 1.0

State-of-the-art set of difficult terminal-based tasks

ModelsModels with a result.45
Top resultBest result on this test.63.8%GPT-5.2
Top-3 spreadPoints from first to third.3.8 pts
UpdatedDate the source changed the results.12 Jan 2026

Result and cost

020406080100$0.003$0.01$0.03$0.1$0.3$1$3$10RESULTUS DOLLARS PER TASK · LOG SCALE

Results

46 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01GPT-5.263.8%$1.424 min
02Claude Sonnet 4.5 Thinking61.3%$0.6619 min
03Gemini 3 Flash Preview60.0%$0.1262 min
04GPT-5 Codex58.8%$1.3811 min
05Claude Opus 4.5 Thinking57.5%$0.9615 min
05GPT-5.1-Codex57.5%$0.12116 min
07Claude Opus 4.556.3%$0.7443 min
08GPT-5.1-Codex-Max53.8%$1.0212 min
09Gemini 3 Pro51.3%$0.3188 min
10GLM-4.750.0%$0.2239 min
10DeepSeek V3.250.0%$0.0257 min
10Claude Haiku 4.5 Thinking50.0%$0.266 min
13GPT-548.8%$0.10116 min
14GPT-5.147.5%$0.10211 min
15Claude Sonnet 4 Thinking45.0%$2.0212 min
16Devstral43.8%$0.4258 min
17GLM-4.642.5%$0.2862 min
18MiniMax M2.141.3%$0.0637 min
18Gemini 2.5 Pro41.3%$0.6646 min
18GLM-4.541.3%$0.1973 min
18DeepSeek V3.141.3%$0.39111 min
22DeepSeek V3.2 (Thinking)40.0%$0.0486 min
22Kimi K2 Thinking40.0%$0.3352 min
22Labs Devstral Small40.0%$0.1355 min
25Grok 438.8%$7.3115 min
25Kimi K2 Instruct38.8%$0.3352 min
27Qwen3 Max36.3%$0.4815 min
27Qwen3 Max36.3%$0.7153 min
29GPT-4.133.8%$1.295 min
31GPT-5 mini30.0%$0.36625 min
32Grok 4.1 Fast (Reasoning)28.8%<$0.014 min
32Magistral Medium28.8%$1.2312 min
34Grok 4 Fast (Reasoning)27.5%$0.1882 min
36GPT-OSS 120B22.5%$0.03954 s
37Grok 4.1 Fast Non Reasoning21.3%$0.0272 min
37Mistral Large21.3%$0.0573 min
37Gemini 2.5 Flash Thinking21.3%$0.4484 min
40Grok Code Fast 120.0%$0.23885 s
41Grok 4 Fast Non Reasoning18.8%$0.31165 s
43Magistral Small15.0%$0.22110 min
44DeepSeek-R113.8%$0.362 min
45Command A6.3%$2.55 min
45Jamba Large 1.76.3%$4.0210 min