Qwen3.5 397B

Qwen3.5 397B is a non-reasoning model from Alibaba. 34 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.51.9 ±3.3
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.86/s
Input / 1MUS dollars per 1M input tokens.$0.6
Output / 1MUS dollars per 1M output tokens.$3.6
ContextMaximum tokens in one request.128K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
53.1
CodingCode writing and repair.
55.0
ReasoningLogic problems and puzzles.
46.5
MultimodalTasks with images and text.
49.7
KnowledgeFacts and expert knowledge.
53.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
46.1
MathMath problems.
54.0

Results

34 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
V*Multimodal95.8%56.2Z.AI
MMLU-ReduxKnowledge94.9%53.8Qwen
Harvard-MIT Mathematics Tournament February 2025Math94.8%53.3Qwen
AIME 2026Math93.3%52.9Qwen
C-EvalKnowledge93.0%C-Eval authors
Harvard-MIT Mathematics Tournament November 2025Math92.7%Qwen
Instruction-Following EvalInstruction92.6%50.8Jeffrey Zhou et al.
MathVisionMultimodal88.6%Qwen
Graduate-Level Google-Proof Q&AKnowledge88.4%59.9David Rein et al.
Harvard-MIT Mathematics Tournament February 2026Math87.9%55.9Qwen
Massive Multitask Language Understanding ProfessionalKnowledge87.8%58.9Yubo Wang et al.
Artificial Analysis GPQA DiamondKnowledge86.1%56.8Artificial Analysis
VideoMMMUMultimodal84.7%Qwen
MMLU-ProXMultilingual84.7%MMLU-ProX authors
τ²-Bench Tool-Agent-User EvaluationAgentic83.9%59.0Victor Barres et al.
LiveCodeBench v6Coding83.6%53.1LiveCodeBench maintainers
MMAnswerBenchMath80.9%Qwen
CharXiv ReasoningMultimodal80.8%58.1CharXiv authors
Massive Multi-discipline Multimodal Understanding ProMultimodal79.0%55.9MMMU-Pro authors
Software Engineering Benchmark VerifiedCoding76.2%58.7Carlos E. Jimenez et al.
MCP-TasksAgentic74.2%Qwen
WideResearchAgentic74.0%59.3Qwen
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge70.4%55.8Xiaoxuan Du et al.
AI-NeedleReasoning68.7%Qwen
τ³-Bench Tool-Agent-User EvaluationAgentic68.4%53.3Sierra Research
ScreenSpot ProMultimodal65.6%46.9Kaixin Li et al.
Artificial Analysis Long Context ReasoningReasoning64.3%52.7Artificial Analysis
LongBench v2Reasoning63.2%LongBench v2 authors
BrowseCompAgentic62.0%52.0OpenAI
NOVA-63Multilingual59.1%Qwen
Claw-EvalAgentic56.8%47.7Bowen Ye et al.
Artificial Analysis MMMU-ProMultimodal52.7%31.5Artificial Analysis
QwenClawBenchAgentic51.8%51.3Qwen
Artificial Analysis IFBenchInstruction51.6%41.5Artificial Analysis
SWE-bench ProCoding50.9%53.2Xiang Deng et al.
Gert Labs Composite Game BenchmarkAgentic46.8%54.0Gert Labs
MCP AtlasAgentic46.1%43.0OpenAI
VITA-BenchAgentic43.7%59.2Meituan LongCat Team
DeepPlanningAgentic37.6%DeepPlanning authors
ToolathlonAgentic36.3%51.8OpenAI
Humanity's Last ExamKnowledge28.7%53.1Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge24.5%44.2Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge21.4%49.2Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge19.8%46.0Artificial Analysis
ResearchClawBenchAgentic14.2%InternScience
Critical Physics TasksReasoning0.9%40.3Artificial Analysis

34 benchmarks count, from 34 of 46 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6