Terminal-Bench 2.0 · BenchLM

A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.

  • In index
  • Agentic
ModelsModels with a result.41
Top resultBest result on this test.82.0%GPT-5.5
Top-3 spreadPoints from first to third.6.9 pts
YearYear of release.N/A

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

41 results
#ModelResultThe score from the source.
01GPT-5.582.0%
03GPT-5.475.1%
11MiMo-V2.565.8%
14GLM-5.163.5%
21Muse Spark59.0%
24GLM-556.2%
32Grok 4.2047.1%
37GLM-4.741.0%