Qwen3.7 Plus

Qwen3.7 Plus is a reasoning model from Alibaba. 49 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.58.2 ±2.8
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.13/s
Input / 1MUS dollars per 1M input tokens.$0.32
Output / 1MUS dollars per 1M output tokens.$1.28
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1454 (#39)

The index is a score out of 100. The ± range shows how much it can change.

39,334 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
58.3
CodingCode writing and repair.
60.3
ReasoningLogic problems and puzzles.
50.8
MultimodalTasks with images and text.
63.8
KnowledgeFacts and expert knowledge.
56.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
60.5
MathMath problems.
55.6

Results

49 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Instruction-Following EvalInstruction94.6%55.1Jeffrey Zhou et al.
MMLU-ReduxKnowledge94.5%53.2Qwen
τ²-Bench Tool-Agent-User EvaluationAgentic93.0%65.5Victor Barres et al.
Harvard-MIT Mathematics Tournament February 2026Math92.9%59.7Qwen
MRCRv2Reasoning91.7%OpenAI
OmniDocBench 1.5Multimodal91.4%OpenAI
MathVisionMultimodal90.3%Qwen
Graduate-Level Google-Proof Q&AKnowledge90.3%61.6David Rein et al.
GPQA DiamondKnowledge90.3%61.6David Rein et al.
Artificial Analysis GPQA DiamondKnowledge90.0%60.8Artificial Analysis
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeCoding89.6%65.8Naman Jain et al.
MMMLUKnowledge89.0%OpenAI
MAXIFEMultilingual88.8%Qwen
Massive Multitask Language Understanding ProfessionalKnowledge88.5%60.0Yubo Wang et al.
Video-MME with subtitleMultimodal88.0%Qwen
MLVU mean averageMultimodal87.4%Qwen
RealWorldQAMultimodal86.9%60.0Qwen
IMOAnswerBenchMath86.0%DeepSeek-AI
CharXiv ReasoningMultimodal85.9%63.9CharXiv authors
VideoMMMUMultimodal85.4%Qwen
MMLU-ProXMultilingual85.4%MMLU-ProX authors
PolyMathMultilingual84.0%Qwen
INCLUDEMultilingual83.0%Qwen
GPQA diamondKnowledge81.8%53.8none effortEpoch AI
SimpleVQAMultimodal81.7%76.1Z.AI
AndroidWorldAgentic81.0%Z.AI
Artificial Analysis MMMU-ProMultimodal80.5%65.5Artificial Analysis
OTIS Mock AIME 2024-2025Math80.0%56.3none effortEpoch AI
Instruction Following BenchmarkInstruction79.1%58.1Benchmark authors
Massive Multi-discipline Multimodal Understanding ProMultimodal79.0%55.9MMMU-Pro authors
ScreenSpot ProMultimodal79.0%60.8Kaixin Li et al.
Artificial Analysis IFBenchInstruction78.0%68.2Artificial Analysis
Software Engineering Benchmark VerifiedCoding77.7%59.9Carlos E. Jimenez et al.
OSWorld-VerifiedAgentic73.3%62.0Tianbao Xie et al.
MCP AtlasAgentic73.2%63.3OpenAI
Artificial Analysis Long Context ReasoningReasoning73.0%58.7Artificial Analysis
Berkeley Function Calling Leaderboard v4Agentic72.9%57.9Arcee AI
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge71.4%56.6Xiaoxuan Du et al.
MedXpertQA MultimodalMultimodal71.0%Meta AI
OCRBench V2Multimodal70.7%OCRBench authors
ERQAMultimodal69.8%64.2Qwen
Claw-EvalAgentic62.7%57.0Bowen Ye et al.
DeepPlanningAgentic62.3%DeepPlanning authors
QwenClawBenchAgentic61.8%60.5Qwen
NOVA-63Multilingual58.8%Qwen
SWE-bench ProCoding57.6%59.7Xiang Deng et al.
Artificial Analysis Coding IndexCoding55.9%58.3Artificial Analysis
SkillsBenchCoding54.3%69.0OpenHands11 Sept 2026Vals AI
Terminal-Bench 2.1Agentic52.8%55.221 Sept 2026Vals AI
Scientific Code BenchmarkCoding51.3%61.9Benchmark authors
ODINW13Multimodal51.1%Qwen
Vibe Code Bench v1.1Coding46.4%61.5OpenHands21 Sept 2026Vals AI
Artificial Analysis SciCodeCoding46.1%56.8Artificial Analysis
VITA-BenchAgentic45.6%60.8Meituan LongCat Team
MMSearch-PlusMultimodal41.4%Z.AI
NL2RepoCoding41.1%57.1MiniMax
Artificial Analysis Humanity's Last ExamKnowledge35.6%63.2Artificial Analysis
Humanity's Last ExamKnowledge34.7%58.2Center for AI Safety et al.
FrontierMath-Tiers-1-3-v2-PrivateMath34.4%50.8none effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge25.2%53.9Artificial Analysis
ApexMath22.7%DeepSeek-AI
Artificial Analysis Omniscience AccuracyKnowledge22.5%41.7Artificial Analysis
APEX-Agents-AAAgentic22.4%58.8Artificial Analysis / Mercor
Artificial Analysis Agentic IndexAgentic19.7%53.2Artificial Analysis
Mystery Game PuzzlesReasoning17.0%52.0Epoch AI
Code MigrationCoding12.9%53.121 Sept 2026Vals AI
GDPval-AA normalizedAgentic12.8%45.4Artificial Analysis
Critical Physics TasksReasoning9.1%57.5Artificial Analysis
Chess PuzzlesReasoning9.0%35.2none effortEpoch AI
OSWorld 2.0Agentic2.8%57.3Mengqi Yuan et al.
Agent Arena command recoveryAgentic-4.462.415 Sept 2026LMArena
Agent Arena steerabilityAgentic-5.461.315 Sept 2026LMArena
Agent Arena task outcomeAgentic-7.159.315 Sept 2026LMArena

49 benchmarks count, from 53 of 73 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6