LMArena
Agent Arena task outcome
Estimated effect of the model on whether a user marks an agent session as finished.
ModelsModels with a result.43
Top resultBest result on this test.19.8Claude Fable 5.1
Top-3 spreadPoints from first to third.6.5
UpdatedDate the source changed the results.15 Sept 2026
Result and price
Results
43 results#ModelResultThe score from the source.
02GPT-6 Astramax effort17.7
04Claude Opus 5max effort12.4
05Muse Spark 1.3max effort10.3
06Hy4 preview9.8
08Qwen3.8-Flash-Next8.9
09GLM-5.3max effort8.5
10GLM-5.3-Flash8.2
14Qwen3.8max effort5.8
15GLM-5.2max effort4.9
16Qwen3.8-27B4.3
21Grok 4.51.9
22Muse Spark 1.10.0
23GPT-5.5xhigh effort-0.6
24GPT-5.4high effort-1.6
25DeepSeek V4 Pro 0423-1.7
27Grok 4.6xhigh effort-2.5
29GPT-5.6 Lunaxhigh effort-5.0
30Gemini 3.1 Pro Preview-5.5
31Claude Sonnet 4.6-5.8
32Qwen3.7max effort-6.4
34Qwen3.7 Plus-7.1
35MiMo-V2.5-Pro-9.0
36Hy3-10.6
37MiniMax M3-12.0
38Mistral Medium 3.5-15.1
39Gemini 3.5 Flash-Lite-17.6
40MiniMax M2.7-18.0
41Inkling-19.3
42Solar Pro 4-20.1
43Inkling-Small-20.4