Claude Sonnet 5.5

Claude Sonnet 5.5 is a reasoning model from Anthropic. 41 benchmarks count toward its index, in 7 categories.

  • Proprietary
  • Reasoning
  • Input: text · image · file
  • Released 2026-09
availableShows if the model has enough results for an index.
IndexOverall score out of 100.78.6±3.0Median 50.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.99/sMedian 62/s
First tokenSeconds to the first output token.3.1 sMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.$2batch $1Median $0.6
Output / 1MUS dollars per 1M output tokens.$10batch $5 US dollars per 1M output tokens in a batch.Median $2.5
ContextMaximum tokens in one request.1MMedian 262K
EloLMArena rating and rank.N/AMedian 1427

Details

From OpenRouter
API IDThe model ID on OpenRouter.
anthropic/claude-sonnet-5.5
Max outputMaximum output tokens per request.
128K
Knowledge cutoffLast date of training data.
N/A
ToolsTool calls on OpenRouter.
Yes
JSON outputOutput follows a given JSON schema.
Yes
Cache read / 1MUS dollars per 1M cached tokens.
$0.2
WeightsLink to the published weights.
N/A
RetiresDate OpenRouter removes the model.
N/A

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.12 counted78.45 / 184
CodingCode writing and repair.10 counted80.44 / 165
ReasoningLogic problems and puzzles.6 counted82.36 / 195
MultimodalTasks with images and text.1 counted76.1—
KnowledgeFacts and expert knowledge.6 counted72.812 / 228
MultilingualTasks in many languages.None countedN/A—
InstructionTasks with strict rules in the prompt.1 counted44.7—
MathMath problems.5 counted78.08 / 129

Results

41 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
OTIS Mock AIME 2024-2025Math100.0%1 / 185
Vibe Code Bench v1.1Coding92.4%1 / 107
GPQA diamond · Epoch AIKnowledge95.6% ±1.42 / 203
Code MigrationCoding69.8%1 / 73
LiveBench CodingCoding91.4%1 / 62
GDPval-AA normalizedAgentic67.0%2 / 121
ProofBench v1.1Math100.0%1 / 49
Critical Physics TasksReasoning31.4%5 / 187
SWE-bench ProCoding81.3%2 / 73
Artificial Analysis SciCodeCoding61.0%4 / 103
Terminal-Bench 4.0Agentic64.1%2 / 44
Mystery Game PuzzlesReasoning65.0% ±4.84 / 72
HealthBench ProfessionalKnowledge69.2%1 / 14
LiveBench ReasoningReasoning91.6%6 / 62
Terminal-Bench 4.0.0Agentic61.8% ±1.52 / 20
Agent Arena task outcomeAgentic13.95 / 48
FrontierMath-Tier-4-v2-PrivateMath80.5% ±6.38 / 63
LiveBench MathematicsMath96.1%8 / 62
Furniture AssemblyReasoning75.0% ±5.54 / 29
Toolathlon-VerifiedAgentic77.8%3 / 21
cursorBench40Coding55.5%2 / 13
FrontierSWE v2Coding61.9%3 / 18
Agent Arena steerabilityAgentic6.98 / 48
IOICoding83.1%9 / 41
Artificial Analysis GDP.pdfAgentic25.8%6 / 23
DeepSWEAgentic71.0%12 / 41
OfficeQA ProMultimodal65.6%5 / 13
LiveBench Agentic CodingAgentic56.3%25 / 62
SimpleQA VerifiedKnowledge46.5% ±1.634 / 77
FrontierCode 1.1 MainCoding46.2%7 / 15
LiveBench LanguageKnowledge78.0%37 / 62
LiveBench Instruction FollowingInstruction56.8%58 / 62
LiveBench Data AnalysisReasoning59.5%59 / 62
MedScribeHealthcare91.1%3 / 107
BioMysteryBenchScience81.1%1 / 24
EMBFinance75.7%3 / 70
ProgramBenchCoding6.5%3 / 61
Tax Agent BenchFinance73.4%4 / 66
Terminal-Bench ScienceScience45.7%4 / 38
Legal Research BenchLegal48.1%8 / 73
SAGEEducation51.8%10 / 90
MedCodeHealthcare52.9%12 / 104
MysteryMechanismScience49.1%3 / 24
Finance Agent (v2)Finance58.1%10 / 74
SRE BenchOther30.2%8 / 32
Public Benefits Bench v1.1Public benefits67.2%13 / 47
CyberBench v1.1Other59.6%33 / 44
Global MMLUMultilingual92.1%—
OfficeQAMultimodal76.9%—
HealthBench raw scoreKnowledge69.4%—
Chartography without toolsMultimodal61.6%—
LatchBio SingleCellBenchKnowledge59.1%—

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.aiOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

More from Anthropic

All models
Claude Opus 5.581.7Claude Fable 5.179.4Claude Opus 576.1Claude Fable 576.0Claude Mythos 572.1