Gemini 3 Flash

Gemini 3 Flash is a non-reasoning model from Google. 22 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.55.3 ±3.9
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.213/s
Input / 1MUS dollars per 1M input tokens.$0.5
Output / 1MUS dollars per 1M output tokens.$3
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1467 (#30)

The index is a score out of 100. The ± range shows how much it can change.

30,225 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
48.6
CodingCode writing and repair.
58.0
ReasoningLogic problems and puzzles.
54.6
MultimodalTasks with images and text.
63.1
KnowledgeFacts and expert knowledge.
58.7
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
45.0
MathMath problems.
58.6

Results

22 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math95.6%65.0high effortEpoch AI
Artificial Analysis Global-MMLU-LiteMultilingual92.7%Artificial Analysis
τ²-bench TelecomAgentic91.2%64.3high effort · Sierra2 Mar 2026Sierra Research
GPQA diamondKnowledge89.4%60.8high effortEpoch AI
τ²-bench AirlineAgentic82.5%58.0high effort · Sierra2 Mar 2026Sierra Research
Artificial Analysis GPQA DiamondKnowledge81.2%51.8Artificial Analysis
Artificial Analysis MMMU-ProMultimodal78.6%63.1Artificial Analysis
τ²-bench RetailAgentic76.8%53.9high effort · Sierra30 Apr 2026Sierra Research
SWE-Bench verifiedCoding75.4%58.0Epoch AI
SWE-bench MultilingualMultilingual72.7%mini-SWE-agent20 Feb 2026SWE-bench team
SWE-bench MultilingualCoding72.7%mini-SWE-agent2 Sept 2026SWE-bench team
SimpleQA VerifiedKnowledge66.8%83.4high effortEpoch AI
Gert Labs Composite Game BenchmarkAgentic56.6%62.7Gert Labs
Artificial Analysis Long Context ReasoningReasoning55.3%46.5Artificial Analysis
Artificial Analysis IFBenchInstruction55.1%45.0Artificial Analysis
FrontierMath-Tiers-1-3-v2-PrivateMath51.2%60.3Epoch AI
Claw-EvalAgentic49.2%35.8Bowen Ye et al.
Artificial Analysis Omniscience AccuracyKnowledge45.8%70.5Artificial Analysis
τ²-Bench Tool-Agent-User EvaluationAgentic43.3%29.9Victor Barres et al.
Chess PuzzlesReasoning40.0%75.2high effortEpoch AI
FrontierMath-2025-02-28-PrivateMath35.6%63.9Epoch AI
τ²-bench BankingAgentic27.3%18.4high effort · Sierra4 Aug 2026Sierra Research
Mystery Game PuzzlesReasoning20.0%55.2high effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge17.9%44.8Artificial Analysis
FrontierMath-Tier-4-v2-PrivateMath17.1%56.1Epoch AI
Artificial Analysis Humanity's Last ExamKnowledge15.0%40.8Artificial Analysis
JobBenchAgentic11.4%42.0Yuetai Li et al.
FrontierMath-Tier-4-2025-07-01-PrivateMath4.2%48.0Epoch AI
Critical Physics TasksReasoning1.4%41.3Artificial Analysis

22 benchmarks count, from 26 of 29 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repositorySWE-bench team, collected directlyNo licence stated for the leaderboard. The harness repo is MIT

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Google

Gemini 3.7 Flash72.3Gemini 3.8 Flash72.1Gemini 3.5 Flash67.3Gemini 3.6 Flash66.8Gemini 3.1 Pro64.7