DeepSWE

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

ModelsModels with a result.36
Top resultBest result on this test.75.4%Muse Spark 1.3
Top-3 spreadPoints from first to third.1.2 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

36 results
#ModelResultThe score from the source.
06SWE-273.0%
09Grok 4.771.0%
11GPT-6 Sol68.8%
15Kimi K367.5%
18GLM-5.366.9%
19GPT-6 Luna66.6%
20Grok 4.665.9%
32Grok 4.553.0%