Qwen3.6-35B-A3B

Qwen3.6-35B-A3B is a reasoning model from Alibaba. 41 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.50.6 ±3.0
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.63/s
Input / 1MUS dollars per 1M input tokens.$0.15
Output / 1MUS dollars per 1M output tokens.$1
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
53.0
CodingCode writing and repair.
50.9
ReasoningLogic problems and puzzles.
45.7
MultimodalTasks with images and text.
53.4
KnowledgeFacts and expert knowledge.
49.4
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
54.4
MathMath problems.
49.6

Results

41 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic95.3%67.2Victor Barres et al.
AI2D test splitMultimodal92.7%Qwen
AIME 2026Math92.7%52.5Qwen
RefCOCO averageMultimodal92.0%RefCOCO dataset authors
Harvard-MIT Mathematics Tournament February 2025Math90.7%49.6Qwen
C-EvalKnowledge90.0%C-Eval authors
OmniDocBench 1.5Multimodal89.9%OpenAI
Harvard-MIT Mathematics Tournament November 2025Math89.1%Qwen
Video-MME with subtitleMultimodal86.6%Qwen
MLVU mean averageMultimodal86.2%Qwen
Graduate-Level Google-Proof Q&AKnowledge86.0%57.6David Rein et al.
RealWorldQAMultimodal85.3%56.9Qwen
Massive Multitask Language Understanding ProfessionalKnowledge85.2%54.8Yubo Wang et al.
GPQA diamondKnowledge84.8%56.6none effortEpoch AI
Artificial Analysis GPQA DiamondKnowledge84.1%54.7Artificial Analysis
VideoMMMUMultimodal83.7%Qwen
Harvard-MIT Mathematics Tournament February 2026Math83.6%52.7Qwen
Video-MME without subtitleMultimodal82.5%Qwen
CC-OCRMultimodal81.9%Qwen
Massive Multi-discipline Multimodal UnderstandingMultimodal81.7%50.9MMMU authors
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeCoding80.4%57.4Naman Jain et al.
MMAnswerBenchMath78.9%Qwen
CharXiv ReasoningMultimodal78.0%54.9CharXiv authors
Massive Multi-discipline Multimodal Understanding ProMultimodal75.3%49.9MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal75.0%58.7Artificial Analysis
Software Engineering Benchmark VerifiedCoding73.4%56.4Carlos E. Jimenez et al.
Artificial Analysis Long Context ReasoningReasoning71.7%57.8Artificial Analysis
OTIS Mock AIME 2024-2025Math68.9%50.1none effortEpoch AI
Claw-EvalAgentic68.7%66.4Bowen Ye et al.
τ³-Bench Tool-Agent-User EvaluationAgentic67.2%52.3Sierra Research
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge64.7%51.0Xiaoxuan Du et al.
Artificial Analysis IFBenchInstruction64.4%54.4Artificial Analysis
MCP AtlasAgentic62.8%55.5OpenAI
WideResearchAgentic60.1%44.3Qwen
SimpleVQAMultimodal58.9%49.2Z.AI
QwenClawBenchAgentic52.6%52.0Qwen
ODINW13Multimodal50.8%Qwen
SWE-bench ProCoding49.5%51.9Xiang Deng et al.
Gert Labs Composite Game BenchmarkAgentic42.6%50.4Gert Labs
Artificial Analysis Coding IndexCoding41.9%48.5Artificial Analysis
Artificial Analysis SciCodeCoding36.6%43.7Artificial Analysis
VITA-BenchAgentic35.6%52.4Meituan LongCat Team
NL2RepoCoding29.4%47.8MiniMax
ToolathlonAgentic26.9%43.0OpenAI
DeepPlanningAgentic25.9%DeepPlanning authors
Artificial Analysis Humanity's Last ExamKnowledge22.2%48.6Artificial Analysis
Mystery Game PuzzlesReasoning22.0%57.3none effortEpoch AI
Humanity's Last ExamKnowledge21.4%46.9Center for AI Safety et al.
FrontierMath-Tiers-1-3-v2-PrivateMath20.4%42.9none effortEpoch AI
GDPval-AA normalizedAgentic19.0%50.2Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge18.8%37.1Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge18.2%45.2Artificial Analysis
Artificial Analysis Agentic IndexAgentic15.0%49.4Artificial Analysis
Chess PuzzlesReasoning4.0%28.7none effortEpoch AI
Critical Physics TasksReasoning0.3%39.0Artificial Analysis

41 benchmarks count, from 42 of 55 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attribution

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6