LMArena
Agent Arena steerability
Estimated effect of the model on whether an agent acts on a correction the user gives it.
ModelsModels with a result.43
Top resultBest result on this test.11.1Claude Opus 4.8
Top-3 spreadPoints from first to third.2.7
UpdatedDate the source changed the results.15 Sept 2026
Result and price
Results
43 results#ModelResultThe score from the source.
06GPT-5.5xhigh effort6.6
07Grok 4.55.9
09GLM-5.2max effort4.9
10Grok 4.6xhigh effort4.7
12Qwen3.8max effort3.6
13GPT-5.4high effort2.5
15GLM-5.3max effort1.9
19GLM-5.3-Flash0.1
20Qwen3.8-27B0.1
22GPT-6 Astramax effort-0.5
24Hy4 preview-1.2
26Hy3-1.8
27DeepSeek V4 Pro 0423-2.1
28Qwen3.7max effort-2.7
30Gemini 3.1 Pro Preview-3.6
31MiMo-V2.5-Pro-4.0
32Muse Spark 1.1-4.2
34Qwen3.8-Flash-Next-4.7
35Claude Sonnet 4.6-5.0
36Qwen3.7 Plus-5.4
37Mistral Medium 3.5-6.0
38MiniMax M3-6.6
39MiniMax M2.7-6.6
40Solar Pro 4-10.1
41Inkling-10.5
42Inkling-Small-11.8
43Gemini 3.5 Flash-Lite-12.8