Software Engineering Benchmark Verified

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

ModelsModels with a result.75
Top resultBest result on this test.96.0%Claude Opus 5
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.2024

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

75 results
#ModelResultThe score from the source.
13MiniMax M380.5%
17GPT-5.280.0%
22BTL-478.4%
25GLM-577.8%
28Inkling77.6%
30Muse Spark77.4%
35Grok 4.2076.7%
43GLM-4.773.8%
66GPT-4.154.6%
69o3-mini49.3%