ProgramBench

Can language models rebuild programs from scratch?

ModelsModels with a result.42
Top resultBest result on this test.N/A
Top-3 spreadPoints from first to third.N/A
UpdatedDate the source changed the results.21 Sept 2026

Result and cost

020406080100$0.03$0.1$0.3$1$3$10$30$100RESULTUS DOLLARS PER TASK · LOG SCALE

Results

43 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
Claude Fable 5.17.0%$58.522.4 h
GPT-6 Astra5.5%$11.5727 min
Claude Opus 53.0%$60.292.9 h
Claude Fable 52.0%$75.682.6 h
GPT-5.6 Sol1.5%$15.2929 min
GLM-5.31.5%$21.964.3 h
Gemini 3.8 Flash1.0%$10.9347 min
Claude Opus 4.81.0%$31.271.5 h
GPT-5.6 Terra0.5%$5.3834 min
GPT-5.50.5%$6.9522 min
GPT-5.4high effort0.5%$2.2716 min
GLM-5.20.5%$12.582.1 h
Claude Sonnet 4.60.5%72 min
Inkling-Small0.5%$0.47929 min
Gemini 3.7 Flash0.0%$7.1322 min
Hy4 preview0.0%$12.812.2 h
GPT-5.6 Luna0.0%$0.94342 min
Qwen3.8 Max0.0%$12.365.9 h
GPT-5.40.0%40 min
Claude Sonnet 50.0%$36.541.5 h
Muse Spark 1.10.0%$13.065.7 h
Gemini 3.5 Flash0.0%18 min
Gemini 3.6 Flash0.0%$5.9248 min
Grok 4.50.0%$4.1219 min
GLM-5.3-Flash0.0%$0.7995.5 h
Claude Opus 4.70.0%54 min
DeepSeek V4 Flash 04230.0%$1.295.6 h
Qwen3.8-27B0.0%$17.995.9 h
DeepSeek V4 Pro 04230.0%63 min
Gemini 3.1 Pro Preview0.0%$2.6313 min
Kimi K2.7 Code0.0%$3.1964 min
Gemini 3 Flash Preview0.0%$0.447 min
GLM-5.10.0%66 min
Inkling0.0%$4.1446 min
GPT-5.4 mini0.0%54 min
Qwen3.6 Plus0.0%33 min
Grok 4.30.0%$0.3564 min
Gemini 3.5 Flash-Lite0.0%$0.3377 min
MiniMax M2.70.0%24 min
Laguna M.10.0%82 min
Laguna XS.20.0%26 min