SWE-Marathon

A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.

ModelsModels with a result.4
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Results

4 results
#ModelResultThe score from the source.
GLM-5.342.5%
Kimi K342.0%