xAI
availableShows if the model has enough results for an index.Grok 4.20
Grok 4.20 is a reasoning model from xAI. 16 benchmarks count toward its score, in 6 categories.
IndexOverall score out of 100.52.5 ±5.0
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.58/s
Input / 1MUS dollars per 1M input tokens.$1.25
Output / 1MUS dollars per 1M output tokens.$2.5
ContextMaximum tokens in one request.2M
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change.
CapabilitiesScore per category, out of 100.
Out of 100Results
16 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 92.2% | 63.1 | — | — | Epoch AI |
| GPQA diamond | Knowledge | 89.3% | 60.7 | — | — | Epoch AI |
| GPQA Diamond | Knowledge | 88.5% | 59.9 | — | — | David Rein et al. |
| Software Engineering Benchmark Verified | Coding | 76.7% | 59.1 | — | — | Carlos E. Jimenez et al. |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 75.2% | 49.7 | — | — | MMMU-Pro authors |
| LiveCodeBench Pro | Coding | 74.2% | — | — | — | LiveCodeBench Pro authors |
| MedXpertQA Multimodal | Multimodal | 65.8% | — | — | — | Meta AI |
| DeepSearchQA | Agentic | 62.8% | 47.8 | — | — | Meta AI |
| CharXiv Reasoning | Multimodal | 60.9% | 35.4 | — | — | CharXiv authors |
| SimpleVQA | Multimodal | 57.4% | 47.4 | — | — | Z.AI |
| ERQA | Multimodal | 54.1% | 49.4 | — | — | Qwen |
| SWE-bench Pro | Coding | 51.8% | 54.1 | — | — | Xiang Deng et al. |
| MedXpertQA Text | Knowledge | 50.2% | — | — | — | Meta AI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 44.9% | 56.7 | — | — | Epoch AI |
| Gert Labs Composite Game Benchmark | Agentic | 38.4% | 46.6 | — | — | Gert Labs |
| Humanity's Last Exam without tools | Knowledge | 31.6% | 55.6 | — | — | OpenAI |
| SimpleQA Verified | Knowledge | 30.2% | 49.5 | — | — | Epoch AI |
| Chess Puzzles | Reasoning | 24.0% | 54.6 | — | — | Epoch AI |
| HealthBench Hard | Knowledge | 20.3% | 57.1 | — | — | Meta AI |
| FrontierMath-Tier-4-v2-Private | Math | 17.1% | 56.1 | — | — | Epoch AI |
16 benchmarks count, from 17 of 20 results. A grey row does not count. Too few models took that benchmark.