LiveBench Agentic Coding

Mean score over agentic tasks in live JavaScript, TypeScript and Python repositories, where the model changes a codebase rather than writing one function.

ModelsModels with a result.55
Top resultBest result on this test.77.3%DeepSeek V4.1 Flash Max
Top-3 spreadPoints from first to third.11.2 pts
UpdatedDate the source changed the results.25 Jun 2026

Result and cost

020406080100$0.01$0.03$0.1$0.3$1$3$10$30RESULTUS DOLLARS PER TASK · LOG SCALE

Results

55 results
#ModelResultThe score from the source.CostUS dollars to run one task.
06Qwen3.8 Max64.6%$1.75
06Smaug Agentic64.6%$1.81
10Qwen3.8-Flash-Next61.6%$0.154
11Qwen3.8-27B61.4%$0.947
12Smaug Flash61.1%$0.078
13GLM-5.360.9%$1.6
14Smaug Mini60.8%$1.07
19GPT-6 Astra Max57.3%$2.43
20Grok 4.657.0%$1.79
21GLM-5.3-Flash56.8%$0.111
22Grok 4.556.5%$0.684
23GPT-5.6 Sol Max56.2%$3.4
24GPT-5.6 Terra Max54.9%$1.35
25Union Alpha54.7%$0
31GLM-5.251.8%$1.46
34GPT-5.2high effort50.3%$1.42
35GPT-5.2-Codex49.4%$0.594
40GPT-5.6 Luna Max48.4%$0.494
43Grok Build 0.145.8%$0.258
44Kimi K2.7 Code45.7%$0.395
46Qwen3.7 Max43.6%$0.796
51Qwen3.6 Plus41.4%$1.55
52MiniMax M340.7%$0.424
54Qwen3.6-27B39.3%$0.813
55Grok 4.318.5%$0.167