Bug Hunt Bench

A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.

ModelsModels with a result.4
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
YearYear of release.2026

Results

4 results
#ModelResultThe score from the source.
GPT-5.6 Sol42.0%
Grok 4.627.0%