Claude Opus 4.8

Claude Opus 4.8 is a reasoning model from Anthropic. 58 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.71.1 ±2.6
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.58/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1453 (#40)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

53,446 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
69.7
CodingCode writing and repair.
71.6
ReasoningLogic problems and puzzles.
70.4
MultimodalTasks with images and text.
71.3
KnowledgeFacts and expert knowledge.
69.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
63.0
MathMath problems.
74.0

Results

58 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math98.3%66.5max effortEpoch AI
United States of America Mathematical Olympiad 2026Math96.7%Mathematical Association of America
τ²-Bench Tool-Agent-User EvaluationAgentic94.4%66.5Victor Barres et al.
LiveBench MathematicsMath94.3%73.9max effort25 Jun 2026LiveBench
Graduate-Level Google-Proof Q&AKnowledge93.6%64.7David Rein et al.
GPQA DiamondKnowledge93.6%64.7David Rein et al.
DeepSearchQAAgentic93.1%72.1Meta AI
ARC-AGI-1 (semi-private)Reasoning92.5%70.9max effortARC Prize Foundation
GPQA DiamondKnowledge92.4%63.61 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge92.0%62.8Artificial Analysis
GPQA diamondKnowledge91.0%62.3max effortEpoch AI
CharXiv ReasoningMultimodal89.9%68.5CharXiv authors
MMLU ProKnowledge89.6%61.71 Sept 2026Vals AI
LiveBench ReasoningReasoning89.2%79.9max effort25 Jun 2026LiveBench
Software Engineering Benchmark VerifiedCoding88.6%68.6Carlos E. Jimenez et al.
SWE-benchCoding88.6%68.61 Sept 2026Vals AI
ScreenSpot ProMultimodal87.9%70.1Kaixin Li et al.
LiveCodeBenchCoding87.8%64.21 Sept 2026Vals AI
INCLUDEMultilingual87.6%Qwen
MMMU ProMultimodal86.6%68.31 Sept 2026Vals AI
BrowseCompAgentic84.3%70.6OpenAI
OSWorld-VerifiedAgentic83.4%71.4Tianbao Xie et al.
Vibe Code Bench v1.1Coding82.7%76.6OpenHands21 Sept 2026Vals AI
MCP AtlasAgentic82.2%70.0OpenAI
LiveBench CodingCoding81.8%73.5max effort25 Jun 2026LiveBench
CharXiv Reasoning without toolsMultimodal80.5%CharXiv authors
FrontierMath-Tiers-1-3-v2-PrivateMath80.0%76.5max effortEpoch AI
LiveBench LanguageKnowledge79.7%66.7max effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning77.7%62.0Artificial Analysis
Artificial Analysis Coding IndexCoding74.3%71.3Artificial Analysis
Gert Labs Composite Game BenchmarkAgentic73.0%77.1Gert Labs
ARC-AGI-2 (semi-private)Reasoning72.1%76.1high effortARC Prize Foundation
LiveBench Instruction FollowingInstruction72.0%73.7max effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic71.9%66.521 Sept 2026Vals AI
Terminal-Bench 2.0Agentic70.0%72.74 Jun 2026Vals AI
SWE-bench ProCoding69.2%70.9Xiang Deng et al.
OfficeQA ProMultimodal66.2%78.5OfficeQA Pro authors
LiveBench Data AnalysisReasoning66.0%47.7max effort25 Jun 2026LiveBench
cursorBench32Coding62.3%70.7Benchmark authors
Artificial Analysis IFBenchInstruction62.2%52.2Artificial Analysis
ToolathlonAgentic59.9%73.8OpenAI
SkillsBenchCoding59.2%73.2OpenHands11 Sept 2026Vals AI
cursorBench31Coding58.4%Benchmark authors
Humanity's Last ExamKnowledge57.9%77.8Center for AI Safety et al.
FrontierMath-Tier-4-v2-PrivateMath56.1%74.9max effortEpoch AI
Artificial Analysis SciCodeCoding54.4%68.4Artificial Analysis
SimpleQA VerifiedKnowledge53.0%70.6max effortEpoch AI
LiveBench Agentic CodingAgentic50.5%64.2max effort25 Jun 2026LiveBench
Humanity's Last Exam without toolsKnowledge49.8%71.0OpenAI
Artificial Analysis Omniscience AccuracyKnowledge48.8%74.2Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge48.7%77.4Artificial Analysis
Code MigrationCoding47.3%75.221 Sept 2026Vals AI
FrontierMath-2025-02-28-PrivateMath47.2%74.7max effortEpoch AI
GDPval-AA normalizedAgentic46.9%71.6Artificial Analysis
FrontierCode 1.1 MainCoding46.5%74.9Cognition
Artificial Analysis AnalystAgentAgentic45.0%73.1Artificial Analysis
Artificial Analysis Agentic IndexAgentic42.6%71.9Artificial Analysis
Furniture AssemblyReasoning42.5%69.8max effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge41.8%74.7Artificial Analysis
τ²-bench BankingAgentic39.7%27.3max effort · Sierra4 Aug 2026Sierra Research
OEIS Open LiteMath39.0%high effortEpoch AI
Mystery Game PuzzlesReasoning36.0%72.2max effortEpoch AI
Chess PuzzlesReasoning34.0%67.5max effortEpoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath31.3%77.4max effortEpoch AI
OEIS OpenMath29.9%high effortEpoch AI
EBR-benchReasoning28.6%68.9max effortEpoch AI
ResearchClawBenchAgentic21.1%InternScience
Terminal-Bench 3.0Agentic21.1%70.9Ryan Marten et al.
Critical Physics TasksReasoning20.9%82.2Artificial Analysis
OSWorld 2.0Agentic20.6%65.8Mengqi Yuan et al.
Terminal-Bench 4.0Agentic16.2%70.921 Sept 2026Vals AI
Agent Arena steerabilityAgentic11.179.8high effort15 Sept 2026LMArena
Agent Arena command recoveryAgentic7.575.8high effort15 Sept 2026LMArena
Agent Arena task outcomeAgentic6.374.5high effort15 Sept 2026LMArena
ARC-AGI-3 (semi-private)Reasoning1.5%high effortARC Prize Foundation
ProgramBenchCoding1.0%21 Sept 2026Vals AI

58 benchmarks count, from 67 of 76 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

GPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Sonnet 567.6