Google
availableShows if the model has enough results for an index.Gemma 4 31B
Gemma 4 31B is a reasoning model from Google in the Gemma 4 family. 24 benchmarks count toward its score, in 8 categories.
IndexOverall score out of 100.48.8 ±3.3
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.35/s
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.256K
EloLMArena rating and rank.1442 (#64)
The index is a score out of 100. The ± range shows how much it can change. A free tier is available.
5,894 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
24 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Artificial Analysis GPQA Diamond | Knowledge | 85.7% | 56.4 | — | — | Artificial Analysis |
| Massive Multitask Language Understanding Professional | Knowledge | 85.2% | 54.8 | — | — | Yubo Wang et al. |
| Graduate-Level Google-Proof Q&A | Knowledge | 84.3% | 56.1 | — | — | David Rein et al. |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 76.9% | 52.5 | — | — | MMMU-Pro authors |
| GPQA diamond | Knowledge | 75.8% | 48.2 | minimal effort | — | Epoch AI |
| Artificial Analysis IFBench | Instruction | 75.6% | 65.8 | — | — | Artificial Analysis |
| React Native Evals | Coding | 75.2% | 52.9 | — | — | Callstack |
| Artificial Analysis MMMU-Pro | Multimodal | 73.4% | 56.8 | — | — | Artificial Analysis |
| OTIS Mock AIME 2024-2025 | Math | 73.3% | 52.6 | minimal effort | — | Epoch AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 69.7% | 56.5 | — | — | Artificial Analysis |
| EuroEval French | Multilingual | 67.9% | 87.0 | — | — | EuroEval |
| EuroEval Portuguese | Multilingual | 67.5% | 86.5 | — | — | EuroEval |
| EuroEval Italian | Multilingual | 66.8% | 85.7 | — | — | EuroEval |
| EuroEval Swedish | Multilingual | 63.1% | 81.1 | — | — | EuroEval |
| EuroEval Dutch | Multilingual | 62.7% | 80.5 | — | — | EuroEval |
| EuroEval Polish | Multilingual | 61.6% | 79.2 | — | — | EuroEval |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 59.9% | 41.8 | — | — | Victor Barres et al. |
| EuroEval Spanish | Multilingual | 55.8% | 71.9 | — | — | EuroEval |
| EuroEval German | Multilingual | 55.7% | 71.8 | — | — | EuroEval |
| EuroEval Portuguese | Multilingual | 55.3% | 71.3 | — | — | EuroEval |
| EuroEval Dutch | Multilingual | 51.8% | 67.0 | — | — | EuroEval |
| EuroEval Italian | Multilingual | 50.3% | 65.1 | — | — | EuroEval |
| EuroEval Swedish | Multilingual | 50.3% | 65.1 | — | — | EuroEval |
| EuroEval French | Multilingual | 50.0% | 64.8 | — | — | EuroEval |
| EuroEval Polish | Multilingual | 46.7% | 60.6 | — | — | EuroEval |
| Artificial Analysis SciCode | Coding | 45.5% | 56.0 | — | — | Artificial Analysis |
| Artificial Analysis Coding Index | Coding | 43.4% | 49.6 | — | — | Artificial Analysis |
| EuroEval German | Multilingual | 42.6% | 55.6 | — | — | EuroEval |
| SWE-Rebench | Coding | 41.6% | — | — | — | Nebius |
| EuroEval Spanish | Multilingual | 40.8% | 53.2 | — | — | EuroEval |
| Terminal-Bench 2.0 | Agentic | 39.3% | 50.7 | — | 4 Jun 2026 | Vals AI |
| Gert Labs Composite Game Benchmark | Agentic | 35.3% | 43.9 | — | — | Gert Labs |
| Humanity's Last Exam | Knowledge | 26.5% | 51.2 | — | — | Center for AI Safety et al. |
| Artificial Analysis Humanity's Last Exam | Knowledge | 23.6% | 50.1 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 20.0% | 38.6 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 19.5% | 45.3 | — | — | OpenAI |
| Artificial Analysis Intelligence Index | Knowledge | 19.0% | 46.1 | — | — | Artificial Analysis |
| SimpleQA Verified | Knowledge | 10.4% | 31.2 | — | — | Epoch AI |
| Artificial Analysis Agentic Index | Agentic | 6.7% | 42.6 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 5.3% | 39.6 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 5.0% | 30.0 | minimal effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 1.4% | 41.3 | — | — | Artificial Analysis |
24 benchmarks count, from 41 of 42 results. A grey row does not count. Too few models took that benchmark.