Alibaba
availableShows if the model has enough results for an index.Qwen3.6 Plus
Qwen3.6 Plus is a reasoning model from Alibaba. 56 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.56.1 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.36/s
Input / 1MUS dollars per 1M input tokens.$0.325
Output / 1MUS dollars per 1M output tokens.$1.95
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1437 (#74)
The index is a score out of 100. The ± range shows how much it can change.
45,319 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
56 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 97.7% | 68.9 | — | — | Victor Barres et al. |
| V* | Multimodal | 96.9% | 57.6 | — | — | Z.AI |
| Harvard-MIT Mathematics Tournament February 2025 | Math | 96.7% | 54.9 | — | — | Qwen |
| AIME 2026 | Math | 95.3% | 54.4 | — | — | Qwen |
| Harvard-MIT Mathematics Tournament November 2025 | Math | 94.6% | — | — | — | Qwen |
| AIME | Math | 94.6% | 57.8 | — | 16 Apr 2026 | Vals AI |
| MMLU-Redux | Knowledge | 94.5% | 53.2 | — | — | Qwen |
| Instruction-Following Eval | Instruction | 94.3% | 54.4 | — | — | Jeffrey Zhou et al. |
| OTIS Mock AIME 2024-2025 | Math | 93.3% | 63.7 | — | — | Epoch AI |
| C-Eval | Knowledge | 93.3% | — | — | — | C-Eval authors |
| Graduate-Level Google-Proof Q&A | Knowledge | 90.4% | 61.7 | — | — | David Rein et al. |
| Massive Multitask Language Understanding Professional | Knowledge | 88.5% | 60.0 | — | — | Yubo Wang et al. |
| GPQA diamond | Knowledge | 88.4% | 59.8 | — | — | Epoch AI |
| Artificial Analysis GPQA Diamond | Knowledge | 88.2% | 58.9 | — | — | Artificial Analysis |
| MathVision | Multimodal | 88.0% | — | — | — | Qwen |
| Harvard-MIT Mathematics Tournament February 2026 | Math | 87.8% | 55.8 | — | — | Qwen |
| MMLU Pro | Knowledge | 87.7% | 58.7 | — | 1 Sept 2026 | Vals AI |
| GPQA Diamond | Knowledge | 87.4% | 58.9 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench v6 | Coding | 87.1% | 56.3 | — | — | LiveCodeBench maintainers |
| Massive Multi-discipline Multimodal Understanding | Multimodal | 86.0% | 55.1 | — | — | MMMU authors |
| LiveCodeBench | Coding | 86.0% | 62.5 | — | 1 Sept 2026 | Vals AI |
| MMLU-ProX | Multilingual | 84.7% | — | — | — | MMLU-ProX authors |
| MMMU Pro | Multimodal | 84.2% | 64.3 | — | 1 Sept 2026 | Vals AI |
| VideoMMMU | Multimodal | 84.0% | — | — | — | Qwen |
| MMAnswerBench | Math | 83.8% | — | — | — | Qwen |
| LiveBench Mathematics | Math | 83.7% | 59.7 | — | 25 Jun 2026 | LiveBench |
| CharXiv Reasoning | Multimodal | 81.5% | 58.9 | — | — | CharXiv authors |
| Software Engineering Benchmark Verified | Coding | 78.8% | 60.7 | — | — | Carlos E. Jimenez et al. |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 78.8% | 55.6 | — | — | MMMU-Pro authors |
| Artificial Analysis Long Context Reasoning | Reasoning | 78.3% | 62.4 | — | — | Artificial Analysis |
| LiveBench Coding | Coding | 78.2% | 67.5 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis MMMU-Pro | Multimodal | 78.0% | 62.4 | — | — | Artificial Analysis |
| LiveBench Reasoning | Reasoning | 75.8% | 61.3 | — | 25 Jun 2026 | LiveBench |
| Instruction Following Benchmark | Instruction | 75.8% | 54.3 | — | — | Benchmark authors |
| Artificial Analysis IFBench | Instruction | 75.2% | 65.4 | — | — | Artificial Analysis |
| LiveBench Language | Knowledge | 75.0% | 61.1 | — | 25 Jun 2026 | LiveBench |
| WideResearch | Agentic | 74.3% | 59.7 | — | — | Qwen |
| MCP-Tasks | Agentic | 74.1% | — | — | — | Qwen |
| SWE-bench | Coding | 73.4% | 56.4 | — | 1 Sept 2026 | Vals AI |
| SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines | Knowledge | 71.6% | 56.8 | — | — | Xiaoxuan Du et al. |
| τ³-Bench Tool-Agent-User Evaluation | Agentic | 70.7% | 55.2 | — | — | Sierra Research |
| LiveBench Data Analysis | Reasoning | 69.9% | 53.1 | — | 25 Jun 2026 | LiveBench |
| AI-Needle | Reasoning | 68.3% | — | — | — | Qwen |
| ScreenSpot Pro | Multimodal | 68.2% | 49.6 | — | — | Kaixin Li et al. |
| LongBench v2 | Reasoning | 62.0% | — | — | — | LongBench v2 authors |
| Claw-Eval | Agentic | 58.8% | 50.8 | — | — | Bowen Ye et al. |
| LiveBench Instruction Following | Instruction | 58.3% | 52.3 | — | 25 Jun 2026 | LiveBench |
| NOVA-63 | Multilingual | 57.9% | — | — | — | Qwen |
| SWE-Bench verified | Coding | 57.9% | 43.9 | — | — | Epoch AI |
| QwenClawBench | Agentic | 57.2% | 56.2 | — | — | Qwen |
| SWE-bench Pro | Coding | 56.6% | 58.7 | — | — | Xiang Deng et al. |
| Artificial Analysis Coding Index | Coding | 54.5% | 57.4 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 53.2% | 55.4 | — | 21 Sept 2026 | Vals AI |
| Gert Labs Composite Game Benchmark | Agentic | 50.6% | 57.4 | — | — | Gert Labs |
| MCP Atlas | Agentic | 48.2% | 44.6 | — | — | OpenAI |
| Terminal-Bench 2.0 | Agentic | 44.9% | 54.8 | — | 4 Jun 2026 | Vals AI |
| VITA-Bench | Agentic | 44.3% | 59.7 | — | — | Meituan LongCat Team |
| SimpleQA Verified | Knowledge | 44.1% | 62.4 | — | — | Epoch AI |
| DeepPlanning | Agentic | 41.5% | — | — | — | DeepPlanning authors |
| LiveBench Agentic Coding | Agentic | 41.4% | 55.6 | — | 25 Jun 2026 | LiveBench |
| Toolathlon | Agentic | 39.8% | 55.1 | — | — | OpenAI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 32.3% | 49.6 | none effort | — | Epoch AI |
| Humanity's Last Exam | Knowledge | 28.8% | 53.2 | — | — | Center for AI Safety et al. |
| Artificial Analysis Humanity's Last Exam | Knowledge | 27.8% | 54.7 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 27.0% | 56.2 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 26.4% | 46.5 | — | — | Artificial Analysis |
| FrontierMath-2025-02-28-Private | Math | 26.2% | 55.1 | — | — | Epoch AI |
| Vibe Code Bench v1.1 | Coding | 25.6% | 52.8 | OpenHands | 21 Sept 2026 | Vals AI |
| GDPval-AA normalized | Agentic | 23.8% | 53.8 | — | — | Artificial Analysis |
| ResearchClawBench | Agentic | 18.0% | — | — | — | InternScience |
| Chess Puzzles | Reasoning | 17.0% | 45.5 | — | — | Epoch AI |
| Mystery Game Puzzles | Reasoning | 12.0% | 46.7 | none effort | — | Epoch AI |
| Code Migration | Coding | 11.1% | 52.0 | — | 21 Sept 2026 | Vals AI |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 8.3% | 52.5 | — | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 2.9% | 44.5 | — | — | Artificial Analysis |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
56 benchmarks count, from 63 of 76 results. A grey row does not count. Too few models took that benchmark.