Terminal-Bench 4.0.0

Agentic terminal task suite run by the Terminal-Bench team.

ModelsModels with a result.15
Top resultBest result on this test.58.2%GPT-6 Astra
Top-3 spreadPoints from first to third.6.4 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.3$1$3$10$30$100RESULTUS DOLLARS PER TASK · LOG SCALE

Results

15 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01GPT-6 Astramax effort · Codex58.2% ±1.4$9.947 min
02Fable 5.1max effort · Claude Code57.9% ±1.9$18.9265 min
03Opus 5max effort · Claude Code51.8% ±1.7$18.0980 min
04Fable 5max effort · Claude Code44.5% ±2.0$22.0270 min
05GLM-5.3max effort · Claude Code41.8% ±1.6$8.271.6 h
06Grok 4.7xhigh effort · Grok Build37.6% ±1.8$11.161.6 h
07GPT-5.6 Solmax effort · Codex37.3% ±1.9$7.740 min
08Opus 4.8max effort · Claude Code23.6% ±1.8$19.6486 min
09GPT-5.6 Terramax effort · Codex21.5% ±1.7$5.2542 min
10Grok 4.6high effort · Grok Build20.3% ±1.6$10.8839 min
12GPT-5.6 Lunamax effort · Codex17.3% ±1.5$1.0568 min
13Grok 4.5high effort · Grok Build12.4% ±1.3$6.3553 min
13Sonnet 5max effort · Claude Code12.4% ±1.6$29.11.8 h