LiveBench Coding

Mean score over writing a solution to a fresh competition problem and completing one whose body was removed.

ModelsModels with a result.55
Top resultBest result on this test.89.3%Claude Opus 5.5
Top-3 spreadPoints from first to third.3.3 pts
UpdatedDate the source changed the results.25 Jun 2026

Result and cost

020406080100$0.0003$0.001$0.003$0.01$0.03$0.1$0.3$1RESULTUS DOLLARS PER TASK · LOG SCALE

Results

55 results
#ModelResultThe score from the source.CostUS dollars to run one task.
04GPT-5.6 Sol Max83.9%$0.107
05GPT-5.2-Codex83.6%$0.029
06GPT-5.6 Luna Max82.9%$0.052
07Smaug Agentic82.5%$0.058
08GPT-5.5xhigh effort82.1%$0.154
08Union Alpha82.1%$0
15GPT-6 Astra Max80.4%$0.296
17GLM-5.279.7%$0.051
20GLM-5.379.0%$0.17
20GLM-5.3-Flash79.0%<$0.01
24GPT-5.6 Terra Max78.2%$0.169
25Qwen3.6 Plus78.2%$0.017
29GPT-5.4xhigh effort77.5%$0.267
33Smaug Flash77.1%<$0.01
34Grok 4.676.8%$0.021
35GPT-5.2high effort76.1%$0.029
38Qwen3.8-27B75.7%$0.019
39Smaug Mini75.4%$0.01
40Qwen3.7 Max74.2%$0.027
41Kimi K2.7 Code74.0%$0.014
43Qwen3.8 Max72.9%$0.113
44Qwen3.8-Flash-Next72.6%$0.028
47Qwen3.6-27B71.8%$0.02
51Grok 4.369.9%$0.011
52Grok 4.568.6%$0.021
53MiniMax M368.2%<$0.01
55Grok Build 0.165.4%<$0.01