LiveBench Language

Mean score over Connections word puzzles, unscrambling film plot summaries and repairing typos. English only.

ModelsModels with a result.55
Top resultBest result on this test.90.7%Claude Fable 5
Top-3 spreadPoints from first to third.1.3 pts
UpdatedDate the source changed the results.25 Jun 2026

Result and cost

020406080100$0.001$0.003$0.01$0.03$0.1$0.3$1RESULTUS DOLLARS PER TASK · LOG SCALE

Results

55 results
#ModelResultThe score from the source.CostUS dollars to run one task.
03GPT-6 Astra Max89.4%$0.222
06GPT-5.6 Sol Max87.7%$0.179
09Union Alpha85.9%$0
12Smaug Agentic84.4%$0.139
14Grok 4.683.7%$0.035
16GPT-5.6 Terra Max82.9%$0.173
17Grok 4.582.8%$0.036
25GLM-5.379.9%$0.241
26GPT-5.2high effort79.8%$0.042
27Qwen3.7 Max79.7%$0.037
28Qwen3.8 Max79.7%$0.135
30Smaug Flash79.2%<$0.01
33Kimi K2.7 Code77.9%$0.035
34GLM-5.3-Flash77.3%<$0.01
35Smaug Mini77.0%$0.018
36MiniMax M376.8%$0.011
37GLM-5.276.2%$0.052
40Qwen3.6 Plus75.0%$0.032
43Qwen3.8-Flash-Next74.6%$0.025
44Qwen3.8-27B74.3%$0.032
46GPT-5.2-Codex73.7%$0.071
47Grok 4.373.6%$0.023
49GPT-5.6 Luna Max72.6%$0.097
50Grok Build 0.172.5%<$0.01
54Qwen3.6-27B63.3%$0.048