Anthropic
availableShows if the model has enough results for an index.Claude Fable 5
Claude Fable 5 is a reasoning model from Anthropic in the Claude Fable family. 53 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.77.5 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.34/s
Input / 1MUS dollars per 1M input tokens.$10 batch $5
Output / 1MUS dollars per 1M output tokens.$50 batch $25 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1493 (#7)
The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.
30,057 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
53 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 99.7% | 67.3 | max effort | — | Epoch AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 98.5% | 69.5 | — | — | Victor Barres et al. |
| ARC-AGI-1 (semi-private) | Reasoning | 98.5% | 73.7 | max effort | — | ARC Prize Foundation |
| LiveBench Mathematics | Math | 96.0% | 76.2 | max effort | 25 Jun 2026 | LiveBench |
| Software Engineering Benchmark Verified | Coding | 95.0% | 73.7 | — | — | Carlos E. Jimenez et al. |
| ProofBench v1.1 | Math | 95.0% | 87.9 | — | 21 Sept 2026 | Vals AI |
| SWE-bench | Coding | 95.0% | 73.7 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Harvey LAB-AA | Agentic | 93.6% | 77.6 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 93.2% | 64.3 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis GPQA Diamond | Knowledge | 92.6% | 63.5 | — | — | Artificial Analysis |
| MMLU Pro | Knowledge | 91.5% | 64.7 | — | 1 Sept 2026 | Vals AI |
| LiveBench Language | Knowledge | 90.7% | 79.9 | max effort | 25 Jun 2026 | LiveBench |
| Vibe Code Bench v1.1 | Coding | 90.4% | 79.8 | OpenHands | 21 Sept 2026 | Vals AI |
| FrontierMath-Tier-4-v2-Private | Math | 90.2% | 91.2 | max effort | — | Epoch AI |
| LiveCodeBench | Coding | 89.8% | 65.9 | — | 1 Sept 2026 | Vals AI |
| LiveBench Reasoning | Reasoning | 89.7% | 80.5 | max effort | 25 Jun 2026 | LiveBench |
| VulcanBench v3 | Coding | 89.5% | 76.7 | — | — | VulcanBench contributors |
| MMMU Pro | Multimodal | 89.3% | 72.7 | — | 1 Sept 2026 | Vals AI |
| ARC-AGI-2 (semi-private) | Reasoning | 89.2% | 84.7 | max effort | — | ARC Prize Foundation |
| FrontierMath-Tiers-1-3-v2-Private | Math | 87.0% | 80.4 | max effort | — | Epoch AI |
| LiveBench Coding | Coding | 86.0% | 80.4 | max effort | 25 Jun 2026 | LiveBench |
| GPQA diamond | Knowledge | 85.9% | 57.5 | max effort | — | Epoch AI |
| OSWorld-Verified | Agentic | 85.0% | 72.9 | — | — | Tianbao Xie et al. |
| Artificial Analysis Long Context Reasoning | Reasoning | 82.3% | 65.2 | — | — | Artificial Analysis |
| LiveBench Data Analysis | Reasoning | 80.5% | 67.8 | max effort | 25 Jun 2026 | LiveBench |
| Terminal-Bench 2.1 | Agentic | 80.5% | 71.6 | — | 21 Sept 2026 | Vals AI |
| SWE-bench Pro | Coding | 80.0% | 81.4 | — | — | Xiang Deng et al. |
| Artificial Analysis Coding Index | Coding | 76.5% | 72.9 | — | — | Artificial Analysis |
| LiveBench Instruction Following | Instruction | 75.8% | 79.5 | max effort | 25 Jun 2026 | LiveBench |
| IOI v1 | Coding | 72.3% | 81.0 | — | 9 Aug 2026 | Vals AI |
| SimpleQA Verified | Knowledge | 70.7% | 87.0 | xhigh effort | — | Epoch AI |
| cursorBench31 | Coding | 70.6% | — | — | — | Benchmark authors |
| cursorBench32 | Coding | 70.5% | 78.4 | — | — | Benchmark authors |
| Artificial Analysis Omniscience Accuracy | Knowledge | 65.4% | 94.8 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 64.4% | 94.2 | — | — | Wisedocs and Artificial Analysis |
| MirrorCode | Coding | 63.9% | — | high effort | — | Epoch AI |
| Artificial Analysis IFBench | Instruction | 63.5% | 53.5 | — | — | Artificial Analysis |
| Terminal-Bench Hard | Agentic | 62.9% | — | — | — | Artificial Analysis |
| LiveBench Agentic Coding | Agentic | 62.2% | 75.2 | max effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis SciCode | Coding | 61.0% | 77.5 | — | — | Artificial Analysis |
| OfficeQA Pro | Multimodal | 57.9% | 70.2 | — | — | OfficeQA Pro authors |
| Artificial Analysis Humanity's Last Exam | Knowledge | 55.5% | 84.7 | — | — | Artificial Analysis |
| Code Migration | Coding | 55.1% | 80.2 | — | 21 Sept 2026 | Vals AI |
| GDPval-AA normalized | Agentic | 54.8% | 77.7 | — | — | Artificial Analysis |
| FrontierCode 1.1 Main | Coding | 53.5% | 81.2 | — | — | Cognition |
| Mystery Game Puzzles | Reasoning | 52.0% | 89.2 | max effort | — | Epoch AI |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 51.1% | 77.3 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 51.0% | 78.7 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 49.6% | 84.5 | — | — | Artificial Analysis |
| Artificial Analysis AnalystAgent | Agentic | 48.8% | 75.8 | — | — | Artificial Analysis |
| FrontierSWE v2 | Coding | 47.0% | 80.4 | — | — | Proximal |
| OEIS Open Lite | Math | 44.0% | — | high effort | — | Epoch AI |
| Chess Puzzles | Reasoning | 41.0% | 76.5 | high effort | — | Epoch AI |
| τ²-bench Banking | Agentic | 39.7% | 27.3 | max effort · Sierra | 4 Aug 2026 | Sierra Research |
| EBR-bench | Reasoning | 39.5% | 76.3 | max effort | — | Epoch AI |
| Blueprint-Bench 2 | Multimodal | 38.6% | — | — | — | Google DeepMind |
| Furniture Assembly | Reasoning | 35.8% | 65.0 | max effort | — | Epoch AI |
| Terminal-Bench 3.0 | Agentic | 34.0% | 80.6 | — | — | Ryan Marten et al. |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 34.0% | 79.1 | — | — | NeoCognition |
| Critical Physics Tasks | Reasoning | 28.6% | 95.0 | — | — | Artificial Analysis |
| Terminal-Bench 4.0 | Agentic | 22.7% | 75.5 | — | 21 Sept 2026 | Vals AI |
| Agent Arena steerability | Agentic | 10.6 | 79.3 | high effort | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 9.1 | 77.6 | high effort | 15 Sept 2026 | LMArena |
| Agent Arena task outcome | Agentic | 6.0 | 74.1 | high effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 2.0% | — | — | 21 Sept 2026 | Vals AI |
| FrontierMath-Erdos | Math | 0.0% | — | max effort | — | Epoch AI |
53 benchmarks count, from 59 of 66 results. A grey row does not count. Too few models took that benchmark.