Zhun Wang et al.
ExploitGym
A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
ModelsModels with a result.14
Top-3 spreadPoints from first to third.19.2 pts
YearYear of release.2026
Result and price
Results
14 results#ModelResultThe score from the source.