ProgramBench: Can Language Models Rebuild Programs From Scratch?

A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.

ModelsModels with a result.12
Top resultBest result on this test.93.0%Claude Opus 5
Top-3 spreadPoints from first to third.5.4 pts
YearYear of release.2026

Result and price

020406080100$0.1$0.3$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

12 results
#ModelResultThe score from the source.
05Kimi K377.8%
06GLM-5.263.7%
11GLM-5.319.0%