DeepPlanning

A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

ModelsModels with a result.7
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Result and price

020406080100$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

7 results
#ModelResultThe score from the source.
Qwen3.7 Plus62.3%
Qwen3.6 Plus41.5%
Qwen3.5 397B37.6%
GLM-514.6%
Kimi K2.514.4%