ExploitGym

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

ModelsModels with a result.14
Top resultBest result on this test.42.4%GPT-6 Astra
Top-3 spreadPoints from first to third.19.2 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30$100$300RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

14 results
#ModelResultThe score from the source.
04GPT-6 Sol22.1%
08GLM-5.315.0%
09GPT-5.513.4%
11GPT-6 Luna11.6%
12GPT-5.46.0%