Claude Opus 5

Claude Opus 5 is a reasoning model from Anthropic. 68 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.77.9 ±2.5
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.54/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1505 (#3)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

42,617 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
75.2
CodingCode writing and repair.
78.3
ReasoningLogic problems and puzzles.
81.0
MultimodalTasks with images and text.
74.5
KnowledgeFacts and expert knowledge.
76.4
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
60.8
MathMath problems.
79.0

Results

68 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
ProofBench v1.1Math99.0%89.521 Sept 2026Vals AI
OTIS Mock AIME 2024-2025Math98.9%66.8max effortEpoch AI
ARC-AGI-1 (semi-private)Reasoning97.5%73.2high effortARC Prize Foundation
SWE-benchCoding97.0%75.41 Sept 2026Vals AI
VulcanBench Coding Intelligence Index v1Coding96.4%VulcanBench contributors
Software Engineering Benchmark VerifiedCoding96.0%74.6Carlos E. Jimenez et al.
LiveBench MathematicsMath95.7%75.8max effort25 Jun 2026LiveBench
DeepSearchQAAgentic95.0%73.6Meta AI
Legal Agent Benchmark mean criterion-pass rate — Harvey held-out setAgentic94.1%Harvey AI
GPQA diamondKnowledge93.9%64.9max effortEpoch AI
Legal Agent Benchmark mean criterion-pass rate — Anthropic harnessAgentic93.7%Harvey AI and Anthropic
Multi-Agent BrowseComp — 10-agent team prerelease configurationAgentic93.6%Anthropic
Artificial Analysis Harvey LAB-AAAgentic93.5%77.5Artificial Analysis
GPQA DiamondKnowledge93.4%64.51 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge93.2%64.1Artificial Analysis
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding93.0%85.8John Yang et al.
Global MMLUMultilingual92.5%Singh et al.
Multi-task Indic Language Understanding BenchmarkMultilingual92.1%Verma et al.
IOI v1Coding91.7%92.19 Aug 2026Vals AI
MMLU ProKnowledge91.6%64.91 Sept 2026Vals AI
ArXivMath June 2026 with toolsMath91.3%MathArena and Anthropic
LiveBench ReasoningReasoning91.2%82.7max effort25 Jun 2026LiveBench
BrowseCompAgentic90.8%76.0OpenAI
ArXivMath June 2026 without toolsMath90.8%MathArena and Anthropic
ARC-AGI-2 (semi-private)Reasoning90.4%85.3max effortARC Prize Foundation
BioMysteryBench Human SolvableKnowledge90.1%Anthropic
MMMU ProMultimodal89.9%73.61 Sept 2026Vals AI
INCLUDEMultilingual89.8%Qwen
MCP-Atlas mean claim coverageAgentic89.1%Anthropic
LiveCodeBenchCoding89.0%65.31 Sept 2026Vals AI
LiveBench LanguageKnowledge88.7%77.5max effort25 Jun 2026LiveBench
Data Research and Analysis with Complex OperationsAgentic88.6%Anthropic
Vibe Code Bench v1.1Coding88.4%78.9OpenHands21 Sept 2026Vals AI
Toolathlon Verified Pass@3Agentic87.0%Anthropic
VulcanBench v3Coding87.0%72.7VulcanBench contributors
MCP AtlasAgentic85.8%72.7OpenAI
FrontierMath-Tiers-1-3-v2-PrivateMath85.6%79.6max effortEpoch AI
GDP.pdf mean criteria pass rate with toolsMultimodal85.5%Surge AI and Anthropic
Artificial Analysis MMMU-ProMultimodal84.7%70.6Artificial Analysis
Terminal-Bench 2.1Agentic84.6%74.021 Sept 2026Vals AI
IOICoding84.3%82.421 Sept 2026Vals AI
LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchKnowledge84.2%Jon M. Laurent et al.
GDP.pdf mean criteria pass rate without toolsMultimodal83.4%Surge AI and Anthropic
ProgramBench hidden-test pass rate after episode 1Coding83.0%Yang et al.
Chartography with image and code toolsMultimodal83.0%Surge AI and Anthropic
BenchCAD Vision2Code voxel IoU with toolsMultimodal82.1%Zhang et al. and Anthropic
LiveBench CodingCoding81.4%72.9max effort25 Jun 2026LiveBench
Toolathlon-VerifiedAgentic80.6%77.8Moonshot AI
Artificial Analysis Long Context ReasoningReasoning79.3%63.1Artificial Analysis
SWE-bench ProCoding79.2%80.6Xiang Deng et al.
RiemannBench with toolsMath79.0%Surge AI and Anthropic
Benchling Molecular Biology Protocols UnderstandingKnowledge78.4%Benchling and Anthropic
OfficeQAMultimodal78.1%Databricks and Anthropic
Artificial Analysis Coding IndexCoding78.0%73.9Artificial Analysis
LiveBench Data AnalysisReasoning74.6%59.5max effort25 Jun 2026LiveBench
HealthBench Professional raw scoreKnowledge73.4%Anthropic
FrontierMath-Tier-4-v2-PrivateMath73.2%83.1max effortEpoch AI
Toolathlon Verified Pass cubedAgentic73.1%Anthropic
LatchBio SpatialBench VerifiedKnowledge72.5%LatchBio and Anthropic
OSWorld 2.0Agentic70.6%89.6Mengqi Yuan et al.
cursorBench32Coding70.0%77.9Benchmark authors
DeepSWEAgentic68.8%73.8Datacurve AI
HealthBench raw scoreKnowledge67.1%Anthropic
OfficeQA ProMultimodal66.9%79.2OfficeQA Pro authors
LiveBench Agentic CodingAgentic65.2%78.1max effort25 Jun 2026LiveBench
Humanity's Last Exam with toolsAgentic64.7%77.2DeepSeek-AI
Humanity's Last ExamKnowledge64.7%83.6Center for AI Safety et al.
LiveBench Instruction FollowingInstruction63.8%60.8max effort25 Jun 2026LiveBench
FrontierCode 1.1 ExtendedCoding63.6%Cognition
Anthropic Organic Chemistry V2 evaluationKnowledge61.6%Anthropic
Molecular Biology Protocols TroubleshootingKnowledge61.1%Anthropic
Artificial Analysis Omniscience AccuracyKnowledge60.9%89.2Artificial Analysis
Furniture AssemblyReasoning60.8%82.8max effortEpoch AI
LatchBio SingleCellBenchKnowledge60.6%LatchBio and Anthropic
SkillsBenchCoding60.4%74.2OpenHands11 Sept 2026Vals AI
GDPval-AA normalizedAgentic60.4%82.0Artificial Analysis
RiemannBench without toolsMath60.0%Surge AI and Anthropic
SimpleQA VerifiedKnowledge59.9%77.0max effortEpoch AI
HealthBench ProfessionalKnowledge59.8%Rebecca Soskin Hicks et al.
Mystery Game PuzzlesReasoning59.0%95.0max effortEpoch AI
HealthBench length-adjusted scoreKnowledge57.8%Anthropic
Code MigrationCoding57.5%81.721 Sept 2026Vals AI
Artificial Analysis AutomationBenchAgentic56.6%67.5Artificial Analysis
Artificial Analysis SciCodeCoding56.4%71.1Artificial Analysis
Humanity's Last Exam without toolsKnowledge56.3%76.5OpenAI
Artificial Analysis Agentic IndexAgentic56.2%83.0Artificial Analysis
Medical Long Context Reasoning (MLCR-AA)Reasoning55.6%86.4Wisedocs and Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge54.9%84.1Artificial Analysis
HLE-VerifiedKnowledge54.4%Weiqi Zhai et al.
Artificial Analysis AnalystAgentAgentic53.8%79.3Artificial Analysis
FrontierCode 1.1 MainCoding53.4%81.1Cognition
FrontierSWE v2Coding52.0%83.2Proximal
Artificial Analysis Intelligence IndexKnowledge50.8%85.9Artificial Analysis
BioMysteryBench Human DifficultKnowledge49.4%Anthropic
τ²-bench BankingAgentic48.7%33.8max effort · Sierra4 Aug 2026Sierra Research
ProteinGym HardKnowledge47.7%Anthropic
Artificial Analysis EnterpriseOps-GymAgentic47.5%72.4Artificial Analysis
EBR-benchReasoning45.7%80.6max effortEpoch AI
Terminal-Bench 4.0Agentic45.5%91.521 Sept 2026Vals AI
Terminal-Bench 3.0Agentic42.7%87.1Ryan Marten et al.
Anthropic Protein Design evaluationKnowledge42.5%Anthropic
Artificial Analysis Tau3-BankingAgentic42.1%69.8Artificial Analysis
International Mathematical Olympiad 2026Math42.0%Anthropic
Chess PuzzlesReasoning42.0%77.8max effortEpoch AI
BenchCAD Vision2Code voxel IoU without toolsMultimodal36.6%Zhang et al. and Anthropic
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic36.0%80.2NeoCognition
ARC-AGI-3 (semi-private)Reasoning30.2%high effortARC Prize Foundation
Chartography without toolsMultimodal29.6%Surge AI and Anthropic
Critical Physics TasksReasoning29.1%95.0Artificial Analysis
Vibe Code Bench 1-100Coding28.5%83.2OpenHands16 Sept 2026Vals AI
Bug Hunt BenchCoding27.0%Pawel Huryn
AutomationBenchAgentic26.0%58.5Moonshot AI
Legal Agent Benchmark all-pass rate — Anthropic harnessAgentic23.6%Harvey AI and Anthropic
Toolathlon Verified average assistant turnsAgentic23.5%Anthropic
Artificial Analysis GDP.pdfAgentic21.6%72.8Artificial Analysis
Agent Arena command recoveryAgentic13.182.1max effort15 Sept 2026LMArena
Agent Arena task outcomeAgentic12.481.4max effort15 Sept 2026LMArena
Legal Agent Benchmark all-pass rate — Harvey held-out setAgentic11.7%Harvey AI
Agent Arena steerabilityAgentic6.875.1max effort15 Sept 2026LMArena
ProgramBenchCoding3.0%21 Sept 2026Vals AI

68 benchmarks count, from 74 of 120 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsSierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Muse Spark 1.375.0 · $4.25GPT-5.6 Sol76.7 · $10GPT-6 Sol76.9 · $10

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Fable 577.5Claude Opus 4.871.1Claude Sonnet 567.6