Google
availableShows if the model has enough results for an index.Gemini 3.5 Flash
Gemini 3.5 Flash is a reasoning model from Google. 54 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.67.3 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.132/s
Input / 1MUS dollars per 1M input tokens.$1.5 batch $0.75
Output / 1MUS dollars per 1M output tokens.$9 batch $4.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1482 (#13)
The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.
38,257 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
54 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 95.6% | 65.0 | high effort | — | Epoch AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 95.3% | 67.2 | — | — | Victor Barres et al. |
| GPQA diamond | Knowledge | 92.8% | 63.9 | high effort | — | Epoch AI |
| GPQA Diamond | Knowledge | 92.7% | 63.8 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 92.7% | 63.8 | — | 1 Sept 2026 | Vals AI |
| ARC-AGI-1 (semi-private) | Reasoning | 92.5% | 70.9 | high effort | — | ARC Prize Foundation |
| Graduate-Level Google-Proof Q&A | Knowledge | 92.2% | 63.4 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 92.2% | 63.0 | — | — | Artificial Analysis |
| MMLU Pro | Knowledge | 89.5% | 61.6 | — | 1 Sept 2026 | Vals AI |
| MMMU Pro | Multimodal | 88.3% | 71.0 | — | 1 Sept 2026 | Vals AI |
| LiveBench Mathematics | Math | 88.2% | 65.8 | high effort | 25 Jun 2026 | LiveBench |
| LiveCodeBench | Coding | 87.6% | 64.0 | — | 1 Sept 2026 | Vals AI |
| LiveBench Language | Knowledge | 84.6% | 72.6 | high effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis MMMU-Pro | Multimodal | 84.3% | 70.1 | — | — | Artificial Analysis |
| CharXiv Reasoning | Multimodal | 84.2% | 62.0 | — | — | CharXiv authors |
| MCP Atlas | Agentic | 83.6% | 71.1 | — | — | OpenAI |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 83.6% | 63.4 | — | — | MMMU-Pro authors |
| LiveBench Reasoning | Reasoning | 82.0% | 69.9 | high effort | 25 Jun 2026 | LiveBench |
| SWE-Bench verified | Coding | 79.3% | 61.2 | high effort | — | Epoch AI |
| SWE-bench | Coding | 78.8% | 60.7 | — | 1 Sept 2026 | Vals AI |
| OSWorld-Verified | Agentic | 78.4% | 66.7 | — | — | Tianbao Xie et al. |
| LiveBench Coding | Coding | 78.2% | 67.5 | high effort | 25 Jun 2026 | LiveBench |
| MRCRv2 | Reasoning | 77.3% | — | — | — | OpenAI |
| Instruction Following Benchmark | Instruction | 76.3% | 54.8 | — | — | Benchmark authors |
| Artificial Analysis IFBench | Instruction | 76.3% | 66.5 | — | — | Artificial Analysis |
| LiveBench Instruction Following | Instruction | 75.6% | 79.3 | high effort | 25 Jun 2026 | LiveBench |
| Terminal-Bench 2.1 | Agentic | 74.2% | 67.8 | — | 21 Sept 2026 | Vals AI |
| ARC-AGI-2 (semi-private) | Reasoning | 72.1% | 76.1 | high effort | — | ARC Prize Foundation |
| SWE-bench Verified | Coding | 71.8% | 55.1 | medium effort · mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| Artificial Analysis Coding Index | Coding | 70.1% | 68.4 | — | — | Artificial Analysis |
| Artificial Analysis Long Context Reasoning | Reasoning | 69.3% | 56.2 | — | — | Artificial Analysis |
| Terminal-Bench 2.0 | Agentic | 67.4% | 70.8 | — | 4 Jun 2026 | Vals AI |
| SWE-bench Multilingual | Coding | 67.0% | — | medium effort · mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| SimpleQA Verified | Knowledge | 66.2% | 82.8 | high effort | — | Epoch AI |
| LiveBench Data Analysis | Reasoning | 64.9% | 46.0 | high effort | 25 Jun 2026 | LiveBench |
| FrontierMath-Tiers-1-3-v2-Private | Math | 62.8% | 66.8 | high effort | — | Epoch AI |
| Gert Labs Composite Game Benchmark | Agentic | 61.9% | 67.3 | — | — | Gert Labs |
| Toolathlon | Agentic | 56.5% | 70.7 | — | — | OpenAI |
| SWE-bench Pro | Coding | 55.1% | 57.3 | — | — | Xiang Deng et al. |
| Artificial Analysis SciCode | Coding | 53.9% | 67.7 | — | — | Artificial Analysis |
| Scientific Code Benchmark | Coding | 53.1% | 63.9 | — | — | Benchmark authors |
| SkillsBench | Coding | 52.7% | 67.7 | OpenHands | 11 Sept 2026 | Vals AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 51.9% | 78.1 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 50.2% | 85.2 | — | — | Artificial Analysis |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 50.1% | 75.9 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 50.0% | 88.2 | high effort | — | Epoch AI |
| cursorBench31 | Coding | 49.8% | — | — | — | Benchmark authors |
| LiveBench Agentic Coding | Agentic | 49.0% | 62.8 | high effort | 25 Jun 2026 | LiveBench |
| cursorBench32 | Coding | 48.8% | 58.0 | — | — | Benchmark authors |
| Vibe Code Bench v1.1 | Coding | 48.7% | 62.4 | OpenHands | 21 Sept 2026 | Vals AI |
| APEX-Agents-AA | Agentic | 47.1% | 78.3 | — | — | Artificial Analysis / Mercor |
| Artificial Analysis Humanity's Last Exam | Knowledge | 42.7% | 70.9 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 42.2% | 68.0 | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 40.2% | 62.8 | — | — | Center for AI Safety et al. |
| FrontierMath-2025-02-28-Private | Math | 39.0% | 67.0 | high effort | — | Epoch AI |
| Blueprint-Bench 2 | Multimodal | 33.6% | — | — | — | Google DeepMind |
| Mystery Game Puzzles | Reasoning | 32.0% | 68.0 | high effort | — | Epoch AI |
| ProofBench v1.1 | Math | 31.0% | 62.0 | — | 21 Sept 2026 | Vals AI |
| OEIS Open Lite | Math | 29.0% | — | high effort | — | Epoch AI |
| Artificial Analysis Agentic Index | Agentic | 27.3% | 59.4 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 26.8% | 60.8 | high effort | — | Epoch AI |
| Code Migration | Coding | 26.7% | 62.0 | — | 21 Sept 2026 | Vals AI |
| MRCR 1M | Reasoning | 26.6% | — | — | — | DeepSeek-AI |
| OEIS Open | Math | 22.2% | — | high effort | — | Epoch AI |
| ResearchClawBench | Agentic | 18.0% | — | — | — | InternScience |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 14.6% | 59.3 | high effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 13.1% | 65.8 | — | — | Artificial Analysis |
| EBR-bench | Reasoning | 4.8% | 52.6 | high effort | — | Epoch AI |
| Terminal-Bench 4.0 | Agentic | 4.0% | 62.3 | — | 21 Sept 2026 | Vals AI |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
54 benchmarks count, from 61 of 70 results. A grey row does not count. Too few models took that benchmark.