Terminal-Bench Science

Research workflow tasks contributed by practicing scientists across five domains

ModelsModels with a result.25
Top resultBest result on this test.65.7%GPT-6 Astra
Top-3 spreadPoints from first to third.42.9 pts
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.1$0.3$1$3$10$30$100RESULTUS DOLLARS PER TASK · LOG SCALE

Results

25 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01GPT-6 Astra65.7%$15.81.5 h
02Claude Fable 5.134.3%$25.432.4 h
03Claude Opus 522.9%$31.372.6 h
04Muse Spark 1.3 Max14.3%$9.872.6 h
05Claude Fable 512.9%$34.182.3 h
05GPT-5.6 Sol12.9%$21.363.9 h
07GPT-5.6 Terra10.0%$10.582.8 h
07Claude Opus 4.810.0%$51.622.4 h
09Muse Spark 1.38.6%$7.882.3 h
09Gemini 3.8 Flash8.6%$9.2648 min
11Gemini 3.7 Flash5.7%$13.2268 min
11GPT-5.6 Luna5.7%$0.55890 min
13Grok 4.64.3%$10.72.5 h
13DeepSeek V4.1 Flash4.3%$2.283.5 h
13GLM-5.34.3%$24.914.0 h
13Gemini 3.5 Flash4.3%$6.456 min
17Claude Sonnet 52.9%$28.062.9 h
18Hy4 preview1.4%$2.143.8 h
18Qwen3.8 Max1.4%$14.963.4 h
18Gemini 3.6 Flash1.4%$14.6286 min
18Qwen3.8-27B1.4%$4.283.7 h
22GLM-5.3-Flash0.0%$1.973.9 h
22DeepSeek V4 Flash 04230.0%$0.9011.8 h
22DeepSeek V4 Pro 04230.0%$2.83.2 h
22Mercury 2.50.0%$0.1941 min