Claude Opus 5.5

Claude Opus 5.5 is a reasoning model from Anthropic. 28 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.82.8 ±3.6
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.85/s
Input / 1MUS dollars per 1M input tokens.$4
Output / 1MUS dollars per 1M output tokens.$20
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
80.3
CodingCode writing and repair.
86.3
ReasoningLogic problems and puzzles.
79.2
MultimodalTasks with images and text.
77.1
KnowledgeFacts and expert knowledge.
87.7
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
63.8
MathMath problems.
77.6

Results

28 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath97.1%77.6max effort25 Jun 2026LiveBench
ArXivMath August 2026 with toolsMath96.9%MathArena and Anthropic
BenchCAD Vision2Code voxel IoU with toolsMultimodal96.2%Zhang et al. and Anthropic
Global MMLUMultilingual94.3%Singh et al.
Multi-task Indic Language Understanding BenchmarkMultilingual93.1%Verma et al.
LiveBench ReasoningReasoning92.2%84.0max effort25 Jun 2026LiveBench
Artificial Analysis Harvey LAB-AAAgentic91.2%74.3Artificial Analysis
Legal Agent Benchmark mean criterion-pass rate — Harvey held-out setAgentic91.2%Harvey AI
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding91.2%84.6John Yang et al.
ArXivMath August 2026 without toolsMath91.2%MathArena and Anthropic
SWE-bench ProCoding89.9%91.0Xiang Deng et al.
BioMysteryBench Human SolvableKnowledge89.3%Anthropic
LiveBench CodingCoding89.3%85.8max effort25 Jun 2026LiveBench
Chartography with image and code toolsMultimodal89.0%Surge AI and Anthropic
Artificial Analysis MMMU-ProMultimodal87.7%74.3Artificial Analysis
LiveBench LanguageKnowledge86.3%74.6max effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning84.7%66.8Artificial Analysis
Anthropic de novo protein-binder design evaluationKnowledge82.6%Anthropic
Toolathlon Verified Pass@3Agentic82.4%Anthropic
LiveBench Data AnalysisReasoning80.3%67.5max effort25 Jun 2026LiveBench
OfficeQAMultimodal78.9%Databricks and Anthropic
Toolathlon-VerifiedAgentic77.8%75.4Moonshot AI
HealthBench Professional raw scoreKnowledge77.1%Anthropic
DeepSWEAgentic74.2%77.7Datacurve AI
Molecular Biology Protocols TroubleshootingKnowledge73.7%Anthropic
BenchCAD Vision2Code voxel IoU without toolsMultimodal73.0%Zhang et al. and Anthropic
Toolathlon Verified Pass cubedAgentic72.2%Anthropic
LatchBio SpatialBench VerifiedKnowledge72.0%LatchBio and Anthropic
LiveBench Agentic CodingAgentic71.7%84.2max effort25 Jun 2026LiveBench
Anthropic biomedical-image-analysis evaluationMultimodal71.4%Anthropic
Artificial Analysis AutomationBenchAgentic69.5%84.5Artificial Analysis
Benchling Molecular Biology Protocols UnderstandingKnowledge69.0%Benchling and Anthropic
HealthBench raw scoreKnowledge68.1%Anthropic
Humanity's Last Exam with toolsAgentic67.7%80.1DeepSeek-AI
OfficeQA ProMultimodal67.7%80.0OfficeQA Pro authors
GDPval-AA normalizedAgentic67.3%87.3Artificial Analysis
Artificial Analysis SciCodeCoding66.9%85.7Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge66.2%95.0Artificial Analysis
LiveBench Instruction FollowingInstruction65.7%63.8max effort25 Jun 2026LiveBench
HealthBench ProfessionalKnowledge65.6%Rebecca Soskin Hicks et al.
Chartography without toolsMultimodal64.4%Surge AI and Anthropic
Humanity's Last Exam without toolsKnowledge64.4%83.3OpenAI
FrontierCode 1.1 ExtendedCoding63.6%Cognition
Anthropic medicinal-chemistry evaluationKnowledge63.5%Anthropic
FrontierSWE v2Coding62.3%88.9Proximal
Artificial Analysis Humanity's Last ExamKnowledge61.4%91.1Artificial Analysis
LatchBio SingleCellBenchKnowledge61.2%LatchBio and Anthropic
HealthBench length-adjusted scoreKnowledge60.6%Anthropic
Anthropic Protein Design evaluationKnowledge60.2%Anthropic
Terminal-Bench-Science 0.1Agentic58.7%Terminal-Bench-Science Team
cursorBench40Coding57.8%Benchmark authors
Artificial Analysis Intelligence IndexKnowledge57.6%94.5Artificial Analysis
Anthropic protein-design library-ranking taskKnowledge56.0%Anthropic
FrontierCode 1.1 MainCoding54.4%82.0Cognition
BioMysteryBench Human DifficultKnowledge50.0%Anthropic
OSWorld 2.0Agentic48.7%79.1Mengqi Yuan et al.
AutomationBenchAgentic40.0%82.4Moonshot AI
Axiom Bio morphology-to-molecule matchingKnowledge34.0%Axiom Bio and Anthropic
Critical Physics TasksReasoning31.7%95.0Artificial Analysis
Toolathlon Verified average assistant turnsAgentic26.9%Anthropic
Artificial Analysis GDP.pdfAgentic26.2%77.7Artificial Analysis
Legal Agent Benchmark all-pass rate — Harvey held-out setAgentic8.3%Harvey AI

28 benchmarks count, from 29 of 62 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not results

More from Anthropic

Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1Claude Sonnet 567.6