Terminal-Bench Hard

A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

  • Not in index
  • Agentic
ModelsModels with a result.6
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.N/A

Results

6 results
#ModelResultThe score from the source.
GPT-5.6 Sol65.9%
MiniMax M342.4%