Ryan Marten et al.
Terminal-Bench 3.0
A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.
ModelsModels with a result.13
Top-3 spreadPoints from first to third.8.7 pts
YearYear of release.2026
Result and price
Results
13 results#ModelResultThe score from the source.