Anthropic
availableShows if the model has enough results for an index.Claude Opus 5
Claude Opus 5 is a reasoning model from Anthropic. 68 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.77.9 ±2.5
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.54/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1505 (#3)
The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.
42,617 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
68 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| ProofBench v1.1 | Math | 99.0% | 89.5 | — | 21 Sept 2026 | Vals AI |
| OTIS Mock AIME 2024-2025 | Math | 98.9% | 66.8 | max effort | — | Epoch AI |
| ARC-AGI-1 (semi-private) | Reasoning | 97.5% | 73.2 | high effort | — | ARC Prize Foundation |
| SWE-bench | Coding | 97.0% | 75.4 | — | 1 Sept 2026 | Vals AI |
| VulcanBench Coding Intelligence Index v1 | Coding | 96.4% | — | — | — | VulcanBench contributors |
| Software Engineering Benchmark Verified | Coding | 96.0% | 74.6 | — | — | Carlos E. Jimenez et al. |
| LiveBench Mathematics | Math | 95.7% | 75.8 | max effort | 25 Jun 2026 | LiveBench |
| DeepSearchQA | Agentic | 95.0% | 73.6 | — | — | Meta AI |
| Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set | Agentic | 94.1% | — | — | — | Harvey AI |
| GPQA diamond | Knowledge | 93.9% | 64.9 | max effort | — | Epoch AI |
| Legal Agent Benchmark mean criterion-pass rate — Anthropic harness | Agentic | 93.7% | — | — | — | Harvey AI and Anthropic |
| Multi-Agent BrowseComp — 10-agent team prerelease configuration | Agentic | 93.6% | — | — | — | Anthropic |
| Artificial Analysis Harvey LAB-AA | Agentic | 93.5% | 77.5 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 93.4% | 64.5 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis GPQA Diamond | Knowledge | 93.2% | 64.1 | — | — | Artificial Analysis |
| ProgramBench: Can Language Models Rebuild Programs From Scratch? | Coding | 93.0% | 85.8 | — | — | John Yang et al. |
| Global MMLU | Multilingual | 92.5% | — | — | — | Singh et al. |
| Multi-task Indic Language Understanding Benchmark | Multilingual | 92.1% | — | — | — | Verma et al. |
| IOI v1 | Coding | 91.7% | 92.1 | — | 9 Aug 2026 | Vals AI |
| MMLU Pro | Knowledge | 91.6% | 64.9 | — | 1 Sept 2026 | Vals AI |
| ArXivMath June 2026 with tools | Math | 91.3% | — | — | — | MathArena and Anthropic |
| LiveBench Reasoning | Reasoning | 91.2% | 82.7 | max effort | 25 Jun 2026 | LiveBench |
| BrowseComp | Agentic | 90.8% | 76.0 | — | — | OpenAI |
| ArXivMath June 2026 without tools | Math | 90.8% | — | — | — | MathArena and Anthropic |
| ARC-AGI-2 (semi-private) | Reasoning | 90.4% | 85.3 | max effort | — | ARC Prize Foundation |
| BioMysteryBench Human Solvable | Knowledge | 90.1% | — | — | — | Anthropic |
| MMMU Pro | Multimodal | 89.9% | 73.6 | — | 1 Sept 2026 | Vals AI |
| INCLUDE | Multilingual | 89.8% | — | — | — | Qwen |
| MCP-Atlas mean claim coverage | Agentic | 89.1% | — | — | — | Anthropic |
| LiveCodeBench | Coding | 89.0% | 65.3 | — | 1 Sept 2026 | Vals AI |
| LiveBench Language | Knowledge | 88.7% | 77.5 | max effort | 25 Jun 2026 | LiveBench |
| Data Research and Analysis with Complex Operations | Agentic | 88.6% | — | — | — | Anthropic |
| Vibe Code Bench v1.1 | Coding | 88.4% | 78.9 | OpenHands | 21 Sept 2026 | Vals AI |
| Toolathlon Verified Pass@3 | Agentic | 87.0% | — | — | — | Anthropic |
| VulcanBench v3 | Coding | 87.0% | 72.7 | — | — | VulcanBench contributors |
| MCP Atlas | Agentic | 85.8% | 72.7 | — | — | OpenAI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 85.6% | 79.6 | max effort | — | Epoch AI |
| GDP.pdf mean criteria pass rate with tools | Multimodal | 85.5% | — | — | — | Surge AI and Anthropic |
| Artificial Analysis MMMU-Pro | Multimodal | 84.7% | 70.6 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 84.6% | 74.0 | — | 21 Sept 2026 | Vals AI |
| IOI | Coding | 84.3% | 82.4 | — | 21 Sept 2026 | Vals AI |
| LABBench2: An Improved Benchmark for AI Systems Performing Biology Research | Knowledge | 84.2% | — | — | — | Jon M. Laurent et al. |
| GDP.pdf mean criteria pass rate without tools | Multimodal | 83.4% | — | — | — | Surge AI and Anthropic |
| ProgramBench hidden-test pass rate after episode 1 | Coding | 83.0% | — | — | — | Yang et al. |
| Chartography with image and code tools | Multimodal | 83.0% | — | — | — | Surge AI and Anthropic |
| BenchCAD Vision2Code voxel IoU with tools | Multimodal | 82.1% | — | — | — | Zhang et al. and Anthropic |
| LiveBench Coding | Coding | 81.4% | 72.9 | max effort | 25 Jun 2026 | LiveBench |
| Toolathlon-Verified | Agentic | 80.6% | 77.8 | — | — | Moonshot AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 79.3% | 63.1 | — | — | Artificial Analysis |
| SWE-bench Pro | Coding | 79.2% | 80.6 | — | — | Xiang Deng et al. |
| RiemannBench with tools | Math | 79.0% | — | — | — | Surge AI and Anthropic |
| Benchling Molecular Biology Protocols Understanding | Knowledge | 78.4% | — | — | — | Benchling and Anthropic |
| OfficeQA | Multimodal | 78.1% | — | — | — | Databricks and Anthropic |
| Artificial Analysis Coding Index | Coding | 78.0% | 73.9 | — | — | Artificial Analysis |
| LiveBench Data Analysis | Reasoning | 74.6% | 59.5 | max effort | 25 Jun 2026 | LiveBench |
| HealthBench Professional raw score | Knowledge | 73.4% | — | — | — | Anthropic |
| FrontierMath-Tier-4-v2-Private | Math | 73.2% | 83.1 | max effort | — | Epoch AI |
| Toolathlon Verified Pass cubed | Agentic | 73.1% | — | — | — | Anthropic |
| LatchBio SpatialBench Verified | Knowledge | 72.5% | — | — | — | LatchBio and Anthropic |
| OSWorld 2.0 | Agentic | 70.6% | 89.6 | — | — | Mengqi Yuan et al. |
| cursorBench32 | Coding | 70.0% | 77.9 | — | — | Benchmark authors |
| DeepSWE | Agentic | 68.8% | 73.8 | — | — | Datacurve AI |
| HealthBench raw score | Knowledge | 67.1% | — | — | — | Anthropic |
| OfficeQA Pro | Multimodal | 66.9% | 79.2 | — | — | OfficeQA Pro authors |
| LiveBench Agentic Coding | Agentic | 65.2% | 78.1 | max effort | 25 Jun 2026 | LiveBench |
| Humanity's Last Exam with tools | Agentic | 64.7% | 77.2 | — | — | DeepSeek-AI |
| Humanity's Last Exam | Knowledge | 64.7% | 83.6 | — | — | Center for AI Safety et al. |
| LiveBench Instruction Following | Instruction | 63.8% | 60.8 | max effort | 25 Jun 2026 | LiveBench |
| FrontierCode 1.1 Extended | Coding | 63.6% | — | — | — | Cognition |
| Anthropic Organic Chemistry V2 evaluation | Knowledge | 61.6% | — | — | — | Anthropic |
| Molecular Biology Protocols Troubleshooting | Knowledge | 61.1% | — | — | — | Anthropic |
| Artificial Analysis Omniscience Accuracy | Knowledge | 60.9% | 89.2 | — | — | Artificial Analysis |
| Furniture Assembly | Reasoning | 60.8% | 82.8 | max effort | — | Epoch AI |
| LatchBio SingleCellBench | Knowledge | 60.6% | — | — | — | LatchBio and Anthropic |
| SkillsBench | Coding | 60.4% | 74.2 | OpenHands | 11 Sept 2026 | Vals AI |
| GDPval-AA normalized | Agentic | 60.4% | 82.0 | — | — | Artificial Analysis |
| RiemannBench without tools | Math | 60.0% | — | — | — | Surge AI and Anthropic |
| SimpleQA Verified | Knowledge | 59.9% | 77.0 | max effort | — | Epoch AI |
| HealthBench Professional | Knowledge | 59.8% | — | — | — | Rebecca Soskin Hicks et al. |
| Mystery Game Puzzles | Reasoning | 59.0% | 95.0 | max effort | — | Epoch AI |
| HealthBench length-adjusted score | Knowledge | 57.8% | — | — | — | Anthropic |
| Code Migration | Coding | 57.5% | 81.7 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis AutomationBench | Agentic | 56.6% | 67.5 | — | — | Artificial Analysis |
| Artificial Analysis SciCode | Coding | 56.4% | 71.1 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 56.3% | 76.5 | — | — | OpenAI |
| Artificial Analysis Agentic Index | Agentic | 56.2% | 83.0 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 55.6% | 86.4 | — | — | Wisedocs and Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 54.9% | 84.1 | — | — | Artificial Analysis |
| HLE-Verified | Knowledge | 54.4% | — | — | — | Weiqi Zhai et al. |
| Artificial Analysis AnalystAgent | Agentic | 53.8% | 79.3 | — | — | Artificial Analysis |
| FrontierCode 1.1 Main | Coding | 53.4% | 81.1 | — | — | Cognition |
| FrontierSWE v2 | Coding | 52.0% | 83.2 | — | — | Proximal |
| Artificial Analysis Intelligence Index | Knowledge | 50.8% | 85.9 | — | — | Artificial Analysis |
| BioMysteryBench Human Difficult | Knowledge | 49.4% | — | — | — | Anthropic |
| τ²-bench Banking | Agentic | 48.7% | 33.8 | max effort · Sierra | 4 Aug 2026 | Sierra Research |
| ProteinGym Hard | Knowledge | 47.7% | — | — | — | Anthropic |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 47.5% | 72.4 | — | — | Artificial Analysis |
| EBR-bench | Reasoning | 45.7% | 80.6 | max effort | — | Epoch AI |
| Terminal-Bench 4.0 | Agentic | 45.5% | 91.5 | — | 21 Sept 2026 | Vals AI |
| Terminal-Bench 3.0 | Agentic | 42.7% | 87.1 | — | — | Ryan Marten et al. |
| Anthropic Protein Design evaluation | Knowledge | 42.5% | — | — | — | Anthropic |
| Artificial Analysis Tau3-Banking | Agentic | 42.1% | 69.8 | — | — | Artificial Analysis |
| International Mathematical Olympiad 2026 | Math | 42.0% | — | — | — | Anthropic |
| Chess Puzzles | Reasoning | 42.0% | 77.8 | max effort | — | Epoch AI |
| BenchCAD Vision2Code voxel IoU without tools | Multimodal | 36.6% | — | — | — | Zhang et al. and Anthropic |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 36.0% | 80.2 | — | — | NeoCognition |
| ARC-AGI-3 (semi-private) | Reasoning | 30.2% | — | high effort | — | ARC Prize Foundation |
| Chartography without tools | Multimodal | 29.6% | — | — | — | Surge AI and Anthropic |
| Critical Physics Tasks | Reasoning | 29.1% | 95.0 | — | — | Artificial Analysis |
| Vibe Code Bench 1-100 | Coding | 28.5% | 83.2 | OpenHands | 16 Sept 2026 | Vals AI |
| Bug Hunt Bench | Coding | 27.0% | — | — | — | Pawel Huryn |
| AutomationBench | Agentic | 26.0% | 58.5 | — | — | Moonshot AI |
| Legal Agent Benchmark all-pass rate — Anthropic harness | Agentic | 23.6% | — | — | — | Harvey AI and Anthropic |
| Toolathlon Verified average assistant turns | Agentic | 23.5% | — | — | — | Anthropic |
| Artificial Analysis GDP.pdf | Agentic | 21.6% | 72.8 | — | — | Artificial Analysis |
| Agent Arena command recovery | Agentic | 13.1 | 82.1 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena task outcome | Agentic | 12.4 | 81.4 | max effort | 15 Sept 2026 | LMArena |
| Legal Agent Benchmark all-pass rate — Harvey held-out set | Agentic | 11.7% | — | — | — | Harvey AI |
| Agent Arena steerability | Agentic | 6.8 | 75.1 | max effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 3.0% | — | — | 21 Sept 2026 | Vals AI |
68 benchmarks count, from 74 of 120 results. A grey row does not count. Too few models took that benchmark.