Vibe Code Bench 1-100

Can models extend a working web application across a long sequence of dependent requests?

ModelsModels with a result.18
Top resultBest result on this test.28.5%Claude Opus 5
Top-3 spreadPoints from first to third.0.9 pts
UpdatedDate the source changed the results.16 Sept 2026

Result and cost

020406080100$0.3$1$3$10$30$100$300RESULTUS DOLLARS PER TASK · LOG SCALE

Results

18 results
#ModelResultThe score from the source.CostUS dollars to run one task.TimeTime to run one task.
01Claude Opus 5OpenHands28.5%$41.4982 min
02Claude Fable 5.1OpenHands28.0%$149.662.9 h
03GPT-6 AstraOpenHands27.6%$84.723.4 h
04GPT-5.6 LunaOpenHands22.6%$4.912.0 h
06GPT-5.6 SolOpenHands20.0%$47.432.6 h
07GLM-5.3OpenHands20.0%$8.2486 min
08Gemini 3.8 FlashOpenHands18.8%$18.9166 min
10DeepSeek V4.1 FlashOpenHands16.4%$0.68828 min
11GLM-5.3-FlashOpenHands16.0%$0.46339 min
12GPT-5.6 TerraOpenHands14.8%$33.331.8 h
13Grok 4.6OpenHands14.8%$26.272.2 h
14Claude Sonnet 5OpenHands13.8%$71.154.3 h
15Qwen3.8 MaxOpenHands12.8%$9.681.6 h
16MiniMax M3OpenHands9.2%$3.0326 min
17InklingOpenHands7.3%$550 min