Terminal-Bench 2.0 · BenchLM
A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.
ModelsModels with a result.41
Top-3 spreadPoints from first to third.6.9 pts
YearYear of release.N/A
Result and price
Results
41 results#ModelResultThe score from the source.