CyberBench v1.1

Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and produce source patches that fix them?

ModelsModels with a result.7
Top resultBest result on this test.73.7%DeepSeek V4.1 Flash
Top-3 spreadPoints from first to third.3.3 pts
UpdatedDate the source changed the results.18 Sept 2026

Result and cost

020406080100$0.03$0.1$0.3$1$3$10RESULTUS DOLLARS PER TASK · LOG SCALE

Results

7 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01DeepSeek V4.1 Flash73.7%$0.10517 min
02Muse Spark 1.3 Max72.7%$3.3420 min
03Claude Fable 5.170.4%$4.0322 min
04Grok 4.666.0%$3.0443 min
05Claude Opus 565.4%$2.8723 min
06Gemini 3.8 Flash43.8%$1.3114 min
07GPT-6 Astra41.1%$1.817 min