Claude Sonnet 4.5 Thinking

Claude Sonnet 4.5 Thinking is a reasoning model from Anthropic in the Claude Sonnet 4.5 family. 11 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.56.1 ±6.9
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.N/A
Output / 1MUS dollars per 1M output tokens.N/A
ContextMaximum tokens in one request.200K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
57.9
CodingCode writing and repair.
51.6
ReasoningLogic problems and puzzles.
N/A
MultimodalTasks with images and text.
56.4
KnowledgeFacts and expert knowledge.
55.9
MultilingualTasks in many languages.
83.4
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
54.8

Results

11 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MGSMMultilingual94.3%9 Jan 2026Vals AI
AIMEMath88.2%54.816 Apr 2026Vals AI
MMLU ProKnowledge87.4%58.21 Sept 2026Vals AI
GPQA DiamondKnowledge81.6%53.61 Sept 2026Vals AI
MMMU ProMultimodal79.3%56.41 Sept 2026Vals AI
LiveCodeBenchCoding73.0%50.71 Sept 2026Vals AI
SWE-benchCoding70.0%53.71 Sept 2026Vals AI
EuroEval FrenchMultilingual66.9%85.8thinking effortEuroEval
EuroEval SwedishMultilingual66.0%84.7thinking effortEuroEval
EuroEval DutchMultilingual65.3%83.8thinking effortEuroEval
EuroEval ItalianMultilingual65.1%83.6thinking effortEuroEval
EuroEval PortugueseMultilingual64.8%83.2thinking effortEuroEval
EuroEval PolishMultilingual63.8%82.0thinking effortEuroEval
Terminal-Bench 1.0Agentic61.3%63.412 Jan 2026Vals AI
EuroEval SpanishMultilingual60.4%77.7thinking effortEuroEval
EuroEval GermanMultilingual57.1%73.6thinking effortEuroEval
Terminal-Bench 2.0Agentic41.6%52.34 Jun 2026Vals AI
Vibe Code Bench v1.1Coding22.6%51.6OpenHands21 Sept 2026Vals AI
IOI v1Coding18.3%50.49 Aug 2026Vals AI

11 benchmarks count, from 18 of 19 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repository

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1