Carlos E. Jimenez et al.
Software Engineering Benchmark Verified
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
ModelsModels with a result.75
Top-3 spreadPoints from first to third.1.0 pts
YearYear of release.2024
Result and price
Results
75 results#ModelResultThe score from the source.