Meta
availableShows if the model has enough results for an index.Muse Spark 1.3
Muse Spark 1.3 is a reasoning model from Meta in the Muse Spark family. 37 benchmarks count toward its score, in 6 categories.
IndexOverall score out of 100.75.0 ±4.2
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.66/s
Input / 1MUS dollars per 1M input tokens.$1.25
Output / 1MUS dollars per 1M output tokens.$4.25
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1490 (#10)
The index is a score out of 100. The ± range shows how much it can change.
4,723 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
37 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 99.2% | 67.0 | xhigh effort | — | Epoch AI |
| OpenAI MRCR v2 8-needle 256K-512K | Reasoning | 98.5% | — | — | — | Meta Superintelligence Labs |
| OpenAI MRCR v2 8-needle 512K-1M | Reasoning | 98.1% | — | — | — | Meta Superintelligence Labs |
| LiveBench Mathematics | Math | 95.9% | 76.1 | xhigh effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis GPQA Diamond | Knowledge | 93.5% | 64.4 | — | — | Artificial Analysis |
| LiveBench Reasoning | Reasoning | 89.7% | 80.5 | xhigh effort | 25 Jun 2026 | LiveBench |
| DeepSearchQA | Agentic | 89.4% | 69.1 | — | — | Meta AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 88.8% | 76.5 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 88.8% | 76.5 | — | — | DeepSeek-AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 83.0% | 65.7 | — | — | Artificial Analysis |
| Vibe Code Bench v1.1 | Coding | 82.9% | 76.6 | OpenHands | 21 Sept 2026 | Vals AI |
| LiveBench Language | Knowledge | 82.8% | 70.4 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Coding | Coding | 81.1% | 72.2 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Data Analysis | Reasoning | 79.6% | 66.5 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Instruction Following | Instruction | 78.0% | 83.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis Coding Index | Coding | 75.8% | 72.4 | — | — | Artificial Analysis |
| DeepSWE | Agentic | 75.4% | 78.6 | — | — | Datacurve AI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 74.0% | 73.1 | max effort | — | Epoch AI |
| Terminal-Bench 2.1 | Agentic | 72.3% | 66.7 | — | 21 Sept 2026 | Vals AI |
| OSWorld 2.0 | Agentic | 66.9% | 87.8 | — | — | Mengqi Yuan et al. |
| JobBench | Agentic | 64.9% | 78.7 | — | — | Yuetai Li et al. |
| LiveBench Agentic Coding | Agentic | 64.1% | 77.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| SWE-Atlas Codebase QnA | Coding | 59.4% | — | — | — | Meta Superintelligence Labs and Scale AI |
| Artificial Analysis SciCode | Coding | 58.8% | 74.5 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 58.7% | 80.6 | — | — | Artificial Analysis |
| Artificial Analysis AutomationBench | Agentic | 57.9% | 69.2 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 55.7% | 82.6 | — | — | Artificial Analysis |
| Artificial Analysis Tau3-Banking | Agentic | 50.5% | 81.2 | — | — | Artificial Analysis |
| AutomationBench | Agentic | 49.4% | 95.0 | — | — | Moonshot AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 48.7% | 77.4 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 48.1% | 82.5 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 46.3% | 70.2 | max effort | — | Epoch AI |
| IOI | Coding | 43.9% | 64.8 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 43.6% | 67.8 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 43.3% | 75.6 | — | — | Wisedocs and Artificial Analysis |
| Chess Puzzles | Reasoning | 38.0% | 72.6 | max effort | — | Epoch AI |
| Code Migration | Coding | 27.6% | 62.6 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis GDP.pdf | Agentic | 26.6% | 78.1 | — | — | Artificial Analysis |
| Mystery Game Puzzles | Reasoning | 25.0% | 60.5 | max effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 24.9% | 90.6 | — | — | Artificial Analysis |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 19.0% | 70.7 | — | — | NeoCognition |
| Terminal-Bench 4.0 | Agentic | 15.2% | 70.1 | — | 21 Sept 2026 | Vals AI |
| Agent Arena task outcome | Agentic | 10.3 | 79.0 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 5.9 | 74.0 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | 1.5 | 69.0 | max effort | 15 Sept 2026 | LMArena |
37 benchmarks count, from 42 of 45 results. A grey row does not count. Too few models took that benchmark.