LiveBench Reasoning

Mean score over web-of-lies, zebra puzzles, spatial reasoning and navigation logic, on questions written after the models were trained.

ModelsModels with a result.55
Top resultBest result on this test.92.7%GPT-6 Astra Max
Top-3 spreadPoints from first to third.1.0 pts
UpdatedDate the source changed the results.25 Jun 2026

Result and cost

020406080100$0.001$0.003$0.01$0.03$0.1$0.3$1RESULTUS DOLLARS PER TASK · LOG SCALE

Results

55 results
#ModelResultThe score from the source.CostUS dollars to run one task.
01GPT-6 Astra Max92.7%$0.39
04GPT-5.6 Sol Max91.7%$0.285
06GPT-5.6 Terra Max90.6%$0.234
07Grok 4.690.5%$0.087
08Smaug Agentic90.3%$0.189
10GPT-5.5xhigh effort89.7%$0.289
17Qwen3.8 Max88.2%$0.131
18GPT-5.4xhigh effort88.1%$0.213
21Qwen3.8-Flash-Next87.4%$0.022
23Grok 4.587.2%$0.088
25Smaug Flash86.2%<$0.01
26GLM-5.385.8%$0.265
27GPT-5.6 Luna Max85.6%$0.138
32Qwen3.7 Max83.3%$0.102
33GPT-5.2high effort83.2%$0.114
34Kimi K2.7 Code82.8%$0.053
34Smaug Mini82.8%$0.019
39Union Alpha80.8%$0
41Qwen3.8-27B80.0%$0.016
43GLM-5.278.6%$0.121
46GPT-5.2-Codex77.7%$0.169
47GLM-5.3-Flash77.6%$0.014
49Grok Build 0.176.4%<$0.01
50Qwen3.6 Plus75.8%$0.094
51MiniMax M374.5%$0.019
53Grok 4.370.8%$0.041
54Qwen3.6-27B70.3%$0.133