Google
availableShows if the model has enough results for an index.Gemini 3.1 Pro
Gemini 3.1 Pro is a reasoning model from Google. 37 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.64.7 ±3.2
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.115/s
Input / 1MUS dollars per 1M input tokens.$2
Output / 1MUS dollars per 1M output tokens.$12
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change.
CapabilitiesScore per category, out of 100.
Out of 100Results
37 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 95.6% | 67.4 | — | — | Victor Barres et al. |
| OTIS Mock AIME 2024-2025 | Math | 95.6% | 65.0 | high effort | — | Epoch AI |
| GPQA diamond | Knowledge | 94.4% | 65.4 | high effort | — | Epoch AI |
| GPQA Diamond | Knowledge | 94.3% | 65.3 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 94.1% | 65.0 | — | — | Artificial Analysis |
| Artificial Analysis Global-MMLU-Lite | Multilingual | 93.2% | — | — | — | Artificial Analysis |
| ScreenSpot Pro | Multimodal | 84.4% | 66.4 | — | — | Kaixin Li et al. |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 83.9% | 63.9 | — | — | MMMU-Pro authors |
| LiveCodeBench Pro | Coding | 82.9% | — | — | — | LiveCodeBench Pro authors |
| Artificial Analysis MMMU-Pro | Multimodal | 82.4% | 67.8 | — | — | Artificial Analysis |
| Artificial Analysis Long Context Reasoning | Reasoning | 82.0% | 65.0 | — | — | Artificial Analysis |
| MedXpertQA Multimodal | Multimodal | 81.3% | — | — | — | Meta AI |
| CharXiv Reasoning | Multimodal | 80.2% | 57.4 | — | — | CharXiv authors |
| React Native Evals | Coding | 78.9% | 57.9 | — | — | Callstack |
| Artificial Analysis IFBench | Instruction | 77.1% | 67.3 | — | — | Artificial Analysis |
| SWE-Bench verified | Coding | 75.6% | 58.2 | — | — | Epoch AI |
| SimpleQA Verified | Knowledge | 73.5% | 89.6 | high effort | — | Epoch AI |
| SimpleVQA | Multimodal | 72.4% | 65.1 | — | — | Z.AI |
| MedXpertQA Text | Knowledge | 71.5% | — | — | — | Meta AI |
| DeepSearchQA | Agentic | 69.7% | 53.3 | — | — | Meta AI |
| ERQA | Multimodal | 69.4% | 63.8 | — | — | Qwen |
| Artificial Analysis Coding Index | Coding | 68.8% | 67.5 | — | — | Artificial Analysis |
| FrontierMath-Tiers-1-3-v2-Private | Math | 59.6% | 65.0 | — | — | Epoch AI |
| Artificial Analysis SciCode | Coding | 58.7% | 74.3 | — | — | Artificial Analysis |
| Claw-Eval | Agentic | 57.8% | 49.3 | — | — | Bowen Ye et al. |
| Gert Labs Composite Game Benchmark | Agentic | 56.9% | 62.9 | — | — | Gert Labs |
| Artificial Analysis Omniscience Accuracy | Knowledge | 54.9% | 81.8 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 49.0% | 86.9 | high effort | — | Epoch AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 47.0% | 75.5 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 45.4% | 67.2 | — | — | OpenAI |
| FrontierMath-2025-02-28-Private | Math | 36.9% | 65.0 | — | — | Epoch AI |
| Mystery Game Puzzles | Reasoning | 34.0% | 70.1 | high effort | — | Epoch AI |
| APEX-Agents-AA | Agentic | 32.0% | 66.3 | — | — | Artificial Analysis / Mercor |
| Artificial Analysis Intelligence Index | Knowledge | 29.7% | 59.6 | — | — | Artificial Analysis |
| ZeroBench | Multimodal | 29.0% | — | — | — | Meta AI |
| FrontierMath-Tier-4-v2-Private | Math | 26.8% | 60.8 | — | — | Epoch AI |
| Furniture Assembly | Reasoning | 26.7% | 58.5 | high effort | — | Epoch AI |
| HealthBench Hard | Knowledge | 20.6% | 57.4 | — | — | Meta AI |
| Critical Physics Tasks | Reasoning | 17.7% | 75.5 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 16.7% | 61.6 | — | — | Epoch AI |
| EBR-bench | Reasoning | 14.3% | 59.1 | — | — | Epoch AI |
| GDPval-AA normalized | Agentic | 13.8% | 46.2 | — | — | Artificial Analysis |
| ResearchClawBench | Agentic | 13.3% | — | — | — | InternScience |
| Artificial Analysis Agentic Index | Agentic | 10.3% | 45.6 | — | — | Artificial Analysis |
| MirrorCode | Coding | 8.9% | — | high effort | — | Epoch AI |
37 benchmarks count, from 38 of 45 results. A grey row does not count. Too few models took that benchmark.