Claude Sonnet 4

Claude Sonnet 4 is a model from Anthropic. 15 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.43.2 ±5.2
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.43/s
Input / 1MUS dollars per 1M input tokens.$3
Output / 1MUS dollars per 1M output tokens.$15
ContextMaximum tokens in one request.200K
EloLMArena rating and rank.1339 (#201)

The index is a score out of 100. The ± range shows how much it can change.

38,966 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
49.2
CodingCode writing and repair.
46.6
ReasoningLogic problems and puzzles.
39.4
MultimodalTasks with images and text.
45.1
KnowledgeFacts and expert knowledge.
45.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
39.7

Results

15 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MGSMMultilingual93.0%9 Jan 2026Vals AI
MATH 500Math90.3%42.79 Jan 2026Vals AI
MATH level 5Math84.4%40.1Epoch AI
τ²-bench RetailAgentic80.5%56.6Anthropic30 Apr 2026Sierra Research
MMLU ProKnowledge79.4%45.71 Sept 2026Vals AI
GPQA diamondKnowledge77.0%49.3Epoch AI
SWE-bench VerifiedCoding74.8%57.5Harness AI1 Sept 2026SWE-bench team
MMMU ProMultimodal72.4%45.11 Sept 2026Vals AI
GPQA DiamondKnowledge69.4%42.41 Sept 2026Vals AI
OTIS Mock AIME 2024-2025Math61.1%45.7Epoch AI
τ²-bench AirlineAgentic60.0%41.9Anthropic2 Mar 2026Sierra Research
LiveCodeBenchCoding59.7%38.61 Sept 2026Vals AI
AIMEMath38.5%31.716 Apr 2026Vals AI
SWE-bench MultimodalMultimodal34.4%OpenHands-Versa17 Nov 2025SWE-bench team
ARC-AGI-1 (semi-private)Reasoning23.8%38.5ARC Prize Foundation
IOI v1Coding6.5%43.69 Aug 2026Vals AI
FrontierMath-2025-02-28-PrivateMath4.1%34.5Epoch AI
ARC-AGI-2 (semi-private)Reasoning1.3%40.4ARC Prize Foundation
FrontierMath-Tier-4-2025-07-01-PrivateMath0.0%43.4Epoch AI

15 benchmarks count, from 17 of 19 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repositorySWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1