Gemini 4 Argon

Gemini 4 Argon is a reasoning model from Google in the Gemini 4 family. 22 benchmarks count toward its index, in 5 categories.

  • Proprietary
  • Reasoning
  • Input: text
  • Released 2026-09
availableShows if the model has enough results for an index.
IndexOverall score out of 100.81.9±5.4Median 50.7
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.—Median 62/s
First tokenSeconds to the first output token.N/AMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.$2Median $0.6
Output / 1MUS dollars per 1M output tokens.$10Median $2.5
ContextMaximum tokens in one request.N/AMedian 262K
EloLMArena rating and rank.N/AMedian 1427

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.11 counted82.11 / 184
CodingCode writing and repair.6 counted81.62 / 165
ReasoningLogic problems and puzzles.2 counted77.610 / 195
MultimodalTasks with images and text.None countedN/A—
KnowledgeFacts and expert knowledge.2 counted79.24 / 228
MultilingualTasks in many languages.None countedN/A—
InstructionTasks with strict rules in the prompt.None countedN/A—
MathMath problems.1 counted84.7—

Results

22 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
Vibe Code Bench v1.1Coding91.9%2 / 107
Agent Arena steerabilityAgentic13.51 / 48
IOICoding100.0%1 / 41
DeepSWEAgentic77.9%1 / 41
Code MigrationCoding68.2%2 / 73
Artificial Analysis SciCodeCoding61.8%3 / 103
Agent Arena task outcomeAgentic15.43 / 48
Critical Physics TasksReasoning27.1%13 / 187
CWE-bench v1Agentic68.0%1 / 11
GDPval-AA normalizedAgentic56.3%11 / 121
ProofBench v1.1Math99.0%5 / 49
Terminal-Bench 4.0Agentic57.6%5 / 44
OSWorld 2.0Agentic69.2%3 / 23
PostTrainBench v1.1Coding45.3%2 / 14
AutomationBenchAgentic51.3%6 / 22
FrontierSWE v2Coding55.0%5 / 18
Agents' Last ExamAgentic39.5%6 / 16
Artificial Analysis GDP.pdfAgentic21.8%10 / 23
Finance Agent (v2)Finance65.4%1 / 74
LegalBenchLegal88.3%3 / 147
MedCodeHealthcare58.8%3 / 104
Tax Agent BenchFinance76.2%2 / 66
SAGEEducation53.6%4 / 90
CyberBench v1.1Other77.9%2 / 44
Legal Research BenchLegal54.8%4 / 73
EMBFinance75.2%4 / 70
SRE BenchOther44.3%3 / 32
Public Benefits Bench v1.1Public benefits69.8%5 / 47
ProgramBenchCoding2.5%7 / 61
Terminal-Bench ScienceScience44.3%5 / 38
MedScribeHealthcare87.4%15 / 107
BioMysteryBenchScience76.3%6 / 24
MysteryMechanismScience45.5%6 / 24
CUA-benchOther4.8%7 / 8
Graphwalks BFS 0K-128KReasoning99.7%—
LVBenchMultimodal91.7%—
GraphWalks BFS 256K–1MReasoning84.2%—
Chartography without toolsMultimodal71.6%—

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.aiVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

More from Google

All models
Gemini 3.8 Flash71.0Gemini 3.7 Flash70.7Gemini 3.5 Flash65.6Gemini 3.6 Flash65.5Gemini 3.1 Pro63.9