Terminal-Bench 4.0

Frontier-difficulty terminal tasks across software, science, ML, operations, hardware, security, and media

ModelsModels with a result.29
Top resultBest result on this test.57.1%GPT-6 Astra
Top-3 spreadPoints from first to third.11.6 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.1$0.3$1$3$10$30$100RESULTUS DOLLARS PER TASK · LOG SCALE

Results

29 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01GPT-6 Astra57.1%$8.2142 min
02Claude Fable 5.149.5%$24.542.5 h
03Claude Opus 545.5%$17.992.5 h
04GPT-5.6 Sol27.8%$8.782.3 h
04Muse Spark 1.3 Max27.8%$5.383.1 h
06GPT-5.6 Terra26.3%$5.9187 min
07GLM-5.325.3%$9.453.4 h
08Qwen3.8 Max24.7%$8.273.5 h
09Claude Fable 522.7%$35.662.3 h
10GLM-5.3-Flash19.7%$0.6451.9 h
11Grok 4.617.2%$6.473.3 h
12Claude Opus 4.816.2%$28.312.7 h
13Muse Spark 1.315.2%$6.243.2 h
14Gemini 3.8 Flash13.1%$11.542.4 h
15DeepSeek V4.1 Flash11.6%$1.853.9 h
16DeepSeek V4 Flash 04239.1%$2.213.5 h
17Claude Sonnet 58.1%$21.373.3 h
18Grok 4.56.6%$6.543.0 h
19Gemini 3.7 Flash6.1%$14.072.0 h
20Muse Spark 1.25.6%$10.012.6 h
21Hy4 preview5.1%$2.13.1 h
22GPT-5.6 Luna4.5%$0.4031.7 h
22Gemini 3.6 Flash4.5%$9.622.1 h
24Gemini 3.5 Flash4.0%$7.3987 min
24Qwen3.8-27B4.0%$2.653.1 h
26Inkling-Small1.5%$0.39913 min
27DeepSeek V4 Pro 04231.0%$3.283.6 h
28Inkling0.0%$2.0638 min
28Mercury 2.50.0%$0.2229 min