Claude Sonnet 4.6

Claude Sonnet 4.6 is a non-reasoning model from Anthropic. 42 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.58.1 ±2.5
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.38/s
Input / 1MUS dollars per 1M input tokens.$3 batch $1.5
Output / 1MUS dollars per 1M output tokens.$15 batch $7.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1458 (#35)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

66,208 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
61.5
CodingCode writing and repair.
61.8
ReasoningLogic problems and puzzles.
51.8
MultimodalTasks with images and text.
57.0
KnowledgeFacts and expert knowledge.
56.9
MultilingualTasks in many languages.
79.9
InstructionTasks with strict rules in the prompt.
30.9
MathMath problems.
55.3

Results

42 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge95.0%76.2Xiaoxuan Du et al.
AIMEMath92.3%56.716 Apr 2026Vals AI
Graduate-Level Google-Proof Q&AKnowledge89.9%61.2David Rein et al.
MMLU ProKnowledge87.3%58.21 Sept 2026Vals AI
ARC-AGI-1 (semi-private)Reasoning86.0%67.8max effortARC Prize Foundation
GPQA DiamondKnowledge85.6%57.31 Sept 2026Vals AI
MMMU ProMultimodal83.6%63.41 Sept 2026Vals AI
LiveCodeBenchCoding82.1%58.91 Sept 2026Vals AI
React Native EvalsCoding80.6%60.3Callstack
Artificial Analysis GPQA DiamondKnowledge79.9%50.4Artificial Analysis
Software Engineering Benchmark VerifiedCoding79.6%61.4Carlos E. Jimenez et al.
τ²-Bench Tool-Agent-User EvaluationAgentic79.5%55.9Victor Barres et al.
Massive Multitask Language Understanding ProfessionalKnowledge79.2%45.3Yubo Wang et al.
GPQA diamondKnowledge78.8%51.0max effortEpoch AI
CharXiv ReasoningMultimodal77.4%54.2CharXiv authors
SWE-benchCoding77.4%59.61 Sept 2026Vals AI
SWE-Bench verifiedCoding75.2%57.9Epoch AI
OSWorld-VerifiedAgentic72.1%60.8Tianbao Xie et al.
OTIS Mock AIME 2024-2025Math71.1%51.3max effortEpoch AI
Artificial Analysis MMMU-ProMultimodal70.6%53.4Artificial Analysis
EuroEval FrenchMultilingual68.5%87.8EuroEval
Artificial Analysis Long Context ReasoningReasoning68.3%55.5Artificial Analysis
Claw-EvalAgentic67.8%65.0Bowen Ye et al.
EuroEval PortugueseMultilingual67.0%85.9EuroEval
EuroEval ItalianMultilingual65.3%83.8EuroEval
CyberGymAgentic65.2%60.8Zhun Wang et al.
EuroEval SwedishMultilingual64.5%82.8EuroEval
Gert Labs Composite Game BenchmarkAgentic62.9%68.3Gert Labs
SWE-RebenchCoding60.7%Nebius
EuroEval PolishMultilingual59.9%77.0EuroEval
EuroEval DutchMultilingual59.7%76.8EuroEval
Terminal-Bench 2.0Agentic59.6%65.24 Jun 2026Vals AI
EuroEval SpanishMultilingual58.5%75.4EuroEval
ARC-AGI-2 (semi-private)Reasoning58.3%69.1max effortARC Prize Foundation
Terminal-Bench 2.1Agentic57.3%57.921 Sept 2026Vals AI
EuroEval GermanMultilingual54.8%70.7EuroEval
Vibe Code Bench v1.1Coding51.5%63.6OpenHands21 Sept 2026Vals AI
SkillsBenchCoding49.1%64.6OpenHands11 Sept 2026Vals AI
Humanity's Last ExamKnowledge49.0%70.3Center for AI Safety et al.
cursorBench31Coding48.8%Benchmark authors
Artificial Analysis IFBenchInstruction41.2%30.9Artificial Analysis
Code MigrationCoding39.9%70.521 Sept 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge38.6%61.6Artificial Analysis
JobBenchAgentic36.9%59.5Yuetai Li et al.
SimpleQA VerifiedKnowledge32.8%51.9max effortEpoch AI
FrontierMath-2025-02-28-PrivateMath32.4%60.9Epoch AI
Artificial Analysis Intelligence IndexKnowledge24.7%53.3Artificial Analysis
FrontierCode 1.1 MainCoding24.3%55.0Cognition
Mystery Game PuzzlesReasoning16.0%50.9low effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge13.3%39.0Artificial Analysis
OSWorld 2.0Agentic8.3%59.9Mengqi Yuan et al.
FrontierMath-Tier-4-2025-07-01-PrivateMath8.3%52.5Epoch AI
Agent Arena command recoveryAgentic5.073.015 Sept 2026LMArena
Chess PuzzlesReasoning3.0%27.4max effortEpoch AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic2.0%61.2NeoCognition
Critical Physics TasksReasoning0.9%40.3Artificial Analysis
ProgramBenchCoding0.5%21 Sept 2026Vals AI
Agent Arena steerabilityAgentic-5.061.715 Sept 2026LMArena
Agent Arena task outcomeAgentic-5.860.815 Sept 2026LMArena

42 benchmarks count, from 56 of 59 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useEpoch AI, collected directlyCC BY — free to use and redistribute with attributionEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1