SWE-Bench verified

Evaluated by Epoch AI on their own harness.

ModelsModels with a result.31
Top resultBest result on this test.83.5%Claude Opus 4.7
Top-3 spreadPoints from first to third.4.1 pts
YearYear of release.N/A

Result and price

020406080100$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

31 results
#ModelResultThe score from the source.
02GPT-5.5xhigh effort80.6% ±1.8
04GLM-5.2max effort78.7% ±1.9
06Qwen3.7 Max77.3% ±1.9
07Claude Opus 4.677.2% ±2.0
08GPT-5.4high effort76.9% ±1.9
09Claude Opus 4.576.7% ±1.9
11Gemini 3.1 Pro75.6% ±2.0
12Gemini 3 Flash75.4% ±2.0
13Claude Sonnet 4.675.2% ±2.0
15GLM-5.174.2% ±2.0
16GPT-5.2high effort73.8% ±2.0
16Kimi K2.573.8% ±2.0
18GPT-5high effort73.6% ±2.0
19Claude Opus 4.173.3% ±2.0
20Gemini 3 Pro72.9% ±2.0
21GLM-572.1% ±2.1
22Claude Sonnet 4.571.3% ±2.1
23Claude Opus 470.7% ±2.1
24GPT-5.1high effort66.9% ±2.1
26o3medium effort62.3% ±2.2
27Claude 3.7 Sonnet61.0% ±2.2
28Qwen3.6 Plus57.9% ±2.2
30GPT-4.148.5% ±2.3
31GPT-4o31.0% ±2.1