Qwen3 Max

Qwen3 Max is a reasoning model from Alibaba. 23 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.44.5 ±4.9
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.34/s
Input / 1MUS dollars per 1M input tokens.$0.78
Output / 1MUS dollars per 1M output tokens.$3.9
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.1413 (#124)

The index is a score out of 100. The ± range shows how much it can change.

8,939 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
45.9
CodingCode writing and repair.
46.8
ReasoningLogic problems and puzzles.
37.3
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
48.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
33.9
MathMath problems.
47.0

Results

23 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MATH level 5Math97.1%46.4Epoch AI
MGSMMultilingual92.1%9 Jan 2026Vals AI
MGSMMultilingual91.8%9 Jan 2026Vals AI
MMLU ProKnowledge84.7%53.91 Sept 2026Vals AI
τ²-bench TelecomAgentic84.2%59.2Alibaba2 Mar 2026Sierra Research
MMLU ProKnowledge83.5%52.11 Sept 2026Vals AI
GPQA DiamondKnowledge82.2%54.11 Sept 2026Vals AI
AIMEMath81.0%51.516 Apr 2026Vals AI
LiveCodeBenchCoding78.2%55.41 Sept 2026Vals AI
GPQA DiamondKnowledge77.8%50.01 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge76.4%46.9Artificial Analysis
τ²-Bench Tool-Agent-User EvaluationAgentic74.3%52.1Victor Barres et al.
OTIS Mock AIME 2024-2025Math73.3%52.6Epoch AI
GPQA diamondKnowledge72.6%45.3Epoch AI
τ²-bench RetailAgentic72.2%50.6Alibaba30 Apr 2026Sierra Research
LiveCodeBenchCoding66.9%45.21 Sept 2026Vals AI
AIMEMath60.7%42.016 Apr 2026Vals AI
τ²-bench AirlineAgentic59.5%41.5Alibaba2 Mar 2026Sierra Research
Artificial Analysis Long Context ReasoningReasoning50.0%42.8Artificial Analysis
SimpleQA VerifiedKnowledge48.7%66.7Epoch AI
Artificial Analysis IFBenchInstruction44.1%33.9Artificial Analysis
Gert Labs Composite Game BenchmarkAgentic43.7%51.4Gert Labs
Terminal-Bench 1.0Agentic36.3%42.312 Jan 2026Vals AI
Terminal-Bench 1.0Agentic36.3%42.312 Jan 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge24.4%44.1Artificial Analysis
Terminal-Bench 2.0Agentic22.5%38.74 Jun 2026Vals AI
FrontierMath-Tiers-1-3-v2-PrivateMath18.9%42.1Epoch AI
Artificial Analysis Intelligence IndexKnowledge15.6%41.9Artificial Analysis
IOI v1Coding14.7%48.39 Aug 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge11.9%37.5Artificial Analysis
IOI v1Coding7.8%44.39 Aug 2026Vals AI
Mystery Game PuzzlesReasoning5.0%39.2Epoch AI
Chess PuzzlesReasoning4.0%28.7Epoch AI
Vibe Code Bench v1.1Coding3.5%43.6OpenHands21 Sept 2026Vals AI
Critical Physics TasksReasoning0.0%38.4Artificial Analysis

23 benchmarks count, from 33 of 35 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISierra Research, collected directlyMIT — results are in the licensed repository

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6