Qwen3.6 Plus

Qwen3.6 Plus is a reasoning model from Alibaba. 56 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.56.1 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.36/s
Input / 1MUS dollars per 1M input tokens.$0.325
Output / 1MUS dollars per 1M output tokens.$1.95
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1437 (#74)

The index is a score out of 100. The ± range shows how much it can change.

45,319 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
55.9
CodingCode writing and repair.
57.9
ReasoningLogic problems and puzzles.
51.2
MultimodalTasks with images and text.
57.3
KnowledgeFacts and expert knowledge.
56.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
56.6
MathMath problems.
56.0

Results

56 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic97.7%68.9Victor Barres et al.
V*Multimodal96.9%57.6Z.AI
Harvard-MIT Mathematics Tournament February 2025Math96.7%54.9Qwen
AIME 2026Math95.3%54.4Qwen
Harvard-MIT Mathematics Tournament November 2025Math94.6%Qwen
AIMEMath94.6%57.816 Apr 2026Vals AI
MMLU-ReduxKnowledge94.5%53.2Qwen
Instruction-Following EvalInstruction94.3%54.4Jeffrey Zhou et al.
OTIS Mock AIME 2024-2025Math93.3%63.7Epoch AI
C-EvalKnowledge93.3%C-Eval authors
Graduate-Level Google-Proof Q&AKnowledge90.4%61.7David Rein et al.
Massive Multitask Language Understanding ProfessionalKnowledge88.5%60.0Yubo Wang et al.
GPQA diamondKnowledge88.4%59.8Epoch AI
Artificial Analysis GPQA DiamondKnowledge88.2%58.9Artificial Analysis
MathVisionMultimodal88.0%Qwen
Harvard-MIT Mathematics Tournament February 2026Math87.8%55.8Qwen
MMLU ProKnowledge87.7%58.71 Sept 2026Vals AI
GPQA DiamondKnowledge87.4%58.91 Sept 2026Vals AI
LiveCodeBench v6Coding87.1%56.3LiveCodeBench maintainers
Massive Multi-discipline Multimodal UnderstandingMultimodal86.0%55.1MMMU authors
LiveCodeBenchCoding86.0%62.51 Sept 2026Vals AI
MMLU-ProXMultilingual84.7%MMLU-ProX authors
MMMU ProMultimodal84.2%64.31 Sept 2026Vals AI
VideoMMMUMultimodal84.0%Qwen
MMAnswerBenchMath83.8%Qwen
LiveBench MathematicsMath83.7%59.725 Jun 2026LiveBench
CharXiv ReasoningMultimodal81.5%58.9CharXiv authors
Software Engineering Benchmark VerifiedCoding78.8%60.7Carlos E. Jimenez et al.
Massive Multi-discipline Multimodal Understanding ProMultimodal78.8%55.6MMMU-Pro authors
Artificial Analysis Long Context ReasoningReasoning78.3%62.4Artificial Analysis
LiveBench CodingCoding78.2%67.525 Jun 2026LiveBench
Artificial Analysis MMMU-ProMultimodal78.0%62.4Artificial Analysis
LiveBench ReasoningReasoning75.8%61.325 Jun 2026LiveBench
Instruction Following BenchmarkInstruction75.8%54.3Benchmark authors
Artificial Analysis IFBenchInstruction75.2%65.4Artificial Analysis
LiveBench LanguageKnowledge75.0%61.125 Jun 2026LiveBench
WideResearchAgentic74.3%59.7Qwen
MCP-TasksAgentic74.1%Qwen
SWE-benchCoding73.4%56.41 Sept 2026Vals AI
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge71.6%56.8Xiaoxuan Du et al.
τ³-Bench Tool-Agent-User EvaluationAgentic70.7%55.2Sierra Research
LiveBench Data AnalysisReasoning69.9%53.125 Jun 2026LiveBench
AI-NeedleReasoning68.3%Qwen
ScreenSpot ProMultimodal68.2%49.6Kaixin Li et al.
LongBench v2Reasoning62.0%LongBench v2 authors
Claw-EvalAgentic58.8%50.8Bowen Ye et al.
LiveBench Instruction FollowingInstruction58.3%52.325 Jun 2026LiveBench
NOVA-63Multilingual57.9%Qwen
SWE-Bench verifiedCoding57.9%43.9Epoch AI
QwenClawBenchAgentic57.2%56.2Qwen
SWE-bench ProCoding56.6%58.7Xiang Deng et al.
Artificial Analysis Coding IndexCoding54.5%57.4Artificial Analysis
Terminal-Bench 2.1Agentic53.2%55.421 Sept 2026Vals AI
Gert Labs Composite Game BenchmarkAgentic50.6%57.4Gert Labs
MCP AtlasAgentic48.2%44.6OpenAI
Terminal-Bench 2.0Agentic44.9%54.84 Jun 2026Vals AI
VITA-BenchAgentic44.3%59.7Meituan LongCat Team
SimpleQA VerifiedKnowledge44.1%62.4Epoch AI
DeepPlanningAgentic41.5%DeepPlanning authors
LiveBench Agentic CodingAgentic41.4%55.625 Jun 2026LiveBench
ToolathlonAgentic39.8%55.1OpenAI
FrontierMath-Tiers-1-3-v2-PrivateMath32.3%49.6none effortEpoch AI
Humanity's Last ExamKnowledge28.8%53.2Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge27.8%54.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge27.0%56.2Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge26.4%46.5Artificial Analysis
FrontierMath-2025-02-28-PrivateMath26.2%55.1Epoch AI
Vibe Code Bench v1.1Coding25.6%52.8OpenHands21 Sept 2026Vals AI
GDPval-AA normalizedAgentic23.8%53.8Artificial Analysis
ResearchClawBenchAgentic18.0%InternScience
Chess PuzzlesReasoning17.0%45.5Epoch AI
Mystery Game PuzzlesReasoning12.0%46.7none effortEpoch AI
Code MigrationCoding11.1%52.021 Sept 2026Vals AI
FrontierMath-Tier-4-2025-07-01-PrivateMath8.3%52.5Epoch AI
Critical Physics TasksReasoning2.9%44.5Artificial Analysis
ProgramBenchCoding0.0%21 Sept 2026Vals AI

56 benchmarks count, from 63 of 76 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not results

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6