Gemini 3.5 Flash

Gemini 3.5 Flash is a reasoning model from Google. 54 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.67.3 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.132/s
Input / 1MUS dollars per 1M input tokens.$1.5 batch $0.75
Output / 1MUS dollars per 1M output tokens.$9 batch $4.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1482 (#13)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

38,257 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
68.3
CodingCode writing and repair.
63.6
ReasoningLogic problems and puzzles.
67.0
MultimodalTasks with images and text.
66.4
KnowledgeFacts and expert knowledge.
71.2
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
66.9
MathMath problems.
63.8

Results

54 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math95.6%65.0high effortEpoch AI
τ²-Bench Tool-Agent-User EvaluationAgentic95.3%67.2Victor Barres et al.
GPQA diamondKnowledge92.8%63.9high effortEpoch AI
GPQA DiamondKnowledge92.7%63.8David Rein et al.
GPQA DiamondKnowledge92.7%63.81 Sept 2026Vals AI
ARC-AGI-1 (semi-private)Reasoning92.5%70.9high effortARC Prize Foundation
Graduate-Level Google-Proof Q&AKnowledge92.2%63.4David Rein et al.
Artificial Analysis GPQA DiamondKnowledge92.2%63.0Artificial Analysis
MMLU ProKnowledge89.5%61.61 Sept 2026Vals AI
MMMU ProMultimodal88.3%71.01 Sept 2026Vals AI
LiveBench MathematicsMath88.2%65.8high effort25 Jun 2026LiveBench
LiveCodeBenchCoding87.6%64.01 Sept 2026Vals AI
LiveBench LanguageKnowledge84.6%72.6high effort25 Jun 2026LiveBench
Artificial Analysis MMMU-ProMultimodal84.3%70.1Artificial Analysis
CharXiv ReasoningMultimodal84.2%62.0CharXiv authors
MCP AtlasAgentic83.6%71.1OpenAI
Massive Multi-discipline Multimodal Understanding ProMultimodal83.6%63.4MMMU-Pro authors
LiveBench ReasoningReasoning82.0%69.9high effort25 Jun 2026LiveBench
SWE-Bench verifiedCoding79.3%61.2high effortEpoch AI
SWE-benchCoding78.8%60.71 Sept 2026Vals AI
OSWorld-VerifiedAgentic78.4%66.7Tianbao Xie et al.
LiveBench CodingCoding78.2%67.5high effort25 Jun 2026LiveBench
MRCRv2Reasoning77.3%OpenAI
Instruction Following BenchmarkInstruction76.3%54.8Benchmark authors
Artificial Analysis IFBenchInstruction76.3%66.5Artificial Analysis
LiveBench Instruction FollowingInstruction75.6%79.3high effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic74.2%67.821 Sept 2026Vals AI
ARC-AGI-2 (semi-private)Reasoning72.1%76.1high effortARC Prize Foundation
SWE-bench VerifiedCoding71.8%55.1medium effort · mini-SWE-agent1 Sept 2026SWE-bench team
Artificial Analysis Coding IndexCoding70.1%68.4Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning69.3%56.2Artificial Analysis
Terminal-Bench 2.0Agentic67.4%70.84 Jun 2026Vals AI
SWE-bench MultilingualCoding67.0%medium effort · mini-SWE-agent2 Sept 2026SWE-bench team
SimpleQA VerifiedKnowledge66.2%82.8high effortEpoch AI
LiveBench Data AnalysisReasoning64.9%46.0high effort25 Jun 2026LiveBench
FrontierMath-Tiers-1-3-v2-PrivateMath62.8%66.8high effortEpoch AI
Gert Labs Composite Game BenchmarkAgentic61.9%67.3Gert Labs
ToolathlonAgentic56.5%70.7OpenAI
SWE-bench ProCoding55.1%57.3Xiang Deng et al.
Artificial Analysis SciCodeCoding53.9%67.7Artificial Analysis
Scientific Code BenchmarkCoding53.1%63.9Benchmark authors
SkillsBenchCoding52.7%67.7OpenHands11 Sept 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge51.9%78.1Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge50.2%85.2Artificial Analysis
Artificial Analysis EnterpriseOps-GymAgentic50.1%75.9Artificial Analysis
Chess PuzzlesReasoning50.0%88.2high effortEpoch AI
cursorBench31Coding49.8%Benchmark authors
LiveBench Agentic CodingAgentic49.0%62.8high effort25 Jun 2026LiveBench
cursorBench32Coding48.8%58.0Benchmark authors
Vibe Code Bench v1.1Coding48.7%62.4OpenHands21 Sept 2026Vals AI
APEX-Agents-AAAgentic47.1%78.3Artificial Analysis / Mercor
Artificial Analysis Humanity's Last ExamKnowledge42.7%70.9Artificial Analysis
GDPval-AA normalizedAgentic42.2%68.0Artificial Analysis
Humanity's Last ExamKnowledge40.2%62.8Center for AI Safety et al.
FrontierMath-2025-02-28-PrivateMath39.0%67.0high effortEpoch AI
Blueprint-Bench 2Multimodal33.6%Google DeepMind
Mystery Game PuzzlesReasoning32.0%68.0high effortEpoch AI
ProofBench v1.1Math31.0%62.021 Sept 2026Vals AI
OEIS Open LiteMath29.0%high effortEpoch AI
Artificial Analysis Agentic IndexAgentic27.3%59.4Artificial Analysis
FrontierMath-Tier-4-v2-PrivateMath26.8%60.8high effortEpoch AI
Code MigrationCoding26.7%62.021 Sept 2026Vals AI
MRCR 1MReasoning26.6%DeepSeek-AI
OEIS OpenMath22.2%high effortEpoch AI
ResearchClawBenchAgentic18.0%InternScience
FrontierMath-Tier-4-2025-07-01-PrivateMath14.6%59.3high effortEpoch AI
Critical Physics TasksReasoning13.1%65.8Artificial Analysis
EBR-benchReasoning4.8%52.6high effortEpoch AI
Terminal-Bench 4.0Agentic4.0%62.321 Sept 2026Vals AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI

54 benchmarks count, from 61 of 70 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsSWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be cited

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 073164.6 · $0.28

More from Google

Gemini 3.7 Flash72.3Gemini 3.8 Flash72.1Gemini 3.6 Flash66.8Gemini 3.1 Pro64.7Gemini 3 Pro60.5