Terminal-Bench-Science 0.1

A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.

ModelsModels with a result.3
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Results

3 results
#ModelResultThe score from the source.
GPT-6 Astra64.6%