Terminal-Bench 3.0

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

ModelsModels with a result.13
Top resultBest result on this test.42.7%Claude Opus 5
Top-3 spreadPoints from first to third.8.7 pts
YearYear of release.2026

Result and price

020406080100$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

13 results
#ModelResultThe score from the source.
04Grok 4.626.5%
07Grok 4.515.7%
13GLM-5.24.6%