John Yang et al.
ProgramBench: Can Language Models Rebuild Programs From Scratch?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.
ModelsModels with a result.12
Top-3 spreadPoints from first to third.5.4 pts
YearYear of release.2026
Result and price
Results
12 results#ModelResultThe score from the source.