LiveBench Instruction Following

Mean score over paraphrasing, simplifying, summarising and story generation, each marked on whether the stated constraints were actually met.

ModelsModels with a result.55
Top resultBest result on this test.81.4%Gemini 3.8 Flash
Top-3 spreadPoints from first to third.3.4 pts
UpdatedDate the source changed the results.25 Jun 2026

Result and cost

020406080100$0.0003$0.001$0.003$0.01$0.03$0.1$0.3$1$3RESULTUS DOLLARS PER TASK · LOG SCALE

Results

55 results
#ModelResultThe score from the source.CostUS dollars to run one task.
04Qwen3.8-Flash-Next77.1%$0.018
07GPT-6 Astra Max75.6%$1.24
11Qwen3.8 Max74.1%$0.082
12Qwen3.7 Max74.0%$0.079
13Smaug Mini73.9%$0.014
15Qwen3.8-27B72.7%$0.013
17Grok 4.671.9%$0.044
18GPT-5.6 Sol Max71.8%$0.418
19Grok 4.571.5%$0.047
20Smaug Agentic71.0%$0.105
22GPT-5.5xhigh effort70.7%$0.227
23GPT-5.4xhigh effort70.2%$0.259
27GLM-5.369.3%$0.271
28Smaug Flash69.3%<$0.01
32GPT-5.2-Codex66.4%$0.071
34Grok Build 0.165.2%<$0.01
36GPT-5.6 Terra Max64.6%$0.332
43Grok 4.362.8%$0.027
45GLM-5.262.3%$0.062
46GPT-5.2high effort61.8%$0.102
48GPT-5.6 Luna Max60.1%$0.153
50Union Alpha59.5%$0
51Qwen3.6 Plus58.3%$0.05
52MiniMax M357.5%$0.019
53Kimi K2.7 Code56.3%$0.047
54Qwen3.6-27B53.2%$0.046
55GLM-5.3-Flash52.8%$0.019