Meta
availableShows if the model has enough results for an index.Muse Glimmer 30B
Muse Glimmer 30B is a reasoning model from Meta in the Muse Glimmer family. 25 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.53.5 ±3.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.100/s
Input / 1MUS dollars per 1M input tokens.$0.3 batch $0.175
Output / 1MUS dollars per 1M output tokens.$1.2 batch $0.75 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.131K
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.
CapabilitiesScore per category, out of 100.
Out of 100Results
25 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| AIME 2026 | Math | 94.7% | 54.0 | — | — | Qwen |
| Artificial Analysis GPQA Diamond | Knowledge | 83.5% | 54.1 | — | — | Artificial Analysis |
| Artificial Analysis Long Context Reasoning | Reasoning | 83.3% | 65.9 | — | — | Artificial Analysis |
| CharXiv Reasoning | Multimodal | 78.8% | 55.8 | — | — | CharXiv authors |
| Instruction Following Benchmark | Instruction | 77.0% | 55.6 | — | — | Benchmark authors |
| Software Engineering Benchmark Verified | Coding | 76.0% | 58.5 | — | — | Carlos E. Jimenez et al. |
| OmniDocBench 1.5 | Multimodal | 75.8% | — | — | — | OpenAI |
| MCP Atlas | Agentic | 75.5% | 65.0 | — | — | OpenAI |
| ScreenSpot Pro | Multimodal | 75.4% | 57.1 | — | — | Kaixin Li et al. |
| DeepSearchQA | Agentic | 74.6% | 57.3 | — | — | Meta AI |
| Artificial Analysis MMMU-Pro | Multimodal | 74.3% | 57.9 | — | — | Artificial Analysis |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 74.0% | 47.8 | — | — | MMMU-Pro authors |
| OSWorld-Verified | Agentic | 65.9% | 55.1 | — | — | Tianbao Xie et al. |
| Terminal-Bench 2.1 (provider run) | Agentic | 51.7% | 54.6 | — | — | DeepSeek-AI |
| SWE-bench Pro | Coding | 51.2% | 53.5 | — | — | Xiang Deng et al. |
| Artificial Analysis Coding Index | Coding | 49.0% | 53.5 | — | — | Artificial Analysis |
| Artificial Analysis SciCode | Coding | 44.9% | 55.2 | — | — | Artificial Analysis |
| Scientific Code Benchmark | Coding | 43.6% | 53.7 | — | — | Benchmark authors |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 34.7% | 54.9 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 27.0% | 47.3 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 22.0% | 48.4 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 20.0% | 55.1 | — | — | Wisedocs and Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 17.5% | 44.3 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 13.7% | 46.1 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 10.5% | 45.7 | — | — | Artificial Analysis |
| Critical Physics Tasks | Reasoning | 2.6% | 43.8 | — | — | Artificial Analysis |
25 benchmarks count, from 25 of 26 results. A grey row does not count. Too few models took that benchmark.