SWE-Rebench

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

ModelsModels with a result.13
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Result and price

020406080100$0.3$1$3$10$30RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

13 results
#ModelResultThe score from the source.
GLM-562.8%
GLM-5.162.7%
Qwen3.5-27B58.9%
GLM-4.758.7%
Kimi K2.558.5%
Composer 258.0%
MiniMax M2.751.9%
Gemma 4 31B41.6%