SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

ModelsModels with a result.71
Top resultBest result on this test.89.9%Claude Opus 5.5
Top-3 spreadPoints from first to third.9.6 pts
YearYear of release.2025

Result and price

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

71 results
#ModelResultThe score from the source.
11Grok 4.564.7%
20GLM-5.262.1%
28MiniMax M359.0%
30GPT-5.558.6%
32GLM-5.158.4%
33GPT-5.457.7%
43MiMo-V2.556.1%
45GPT-5.255.6%
47GLM-555.1%
49Inkling54.3%
55Muse Spark52.4%
56Grok 4.2051.8%