GPT-5.2

GPT-5.2 is a reasoning model from OpenAI. 44 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.62.2 ±2.4
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.43/s
Input / 1MUS dollars per 1M input tokens.$1.75 batch $0.875
Output / 1MUS dollars per 1M output tokens.$14 batch $7 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.1412 (#125)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

78,967 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
56.5
CodingCode writing and repair.
62.8
ReasoningLogic problems and puzzles.
67.5
MultimodalTasks with images and text.
51.3
KnowledgeFacts and expert knowledge.
62.2
MultilingualTasks in many languages.
75.4
InstructionTasks with strict rules in the prompt.
61.6
MathMath problems.
65.9

Results

44 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Artificial Analysis AIME 2025Math99.0%Artificial Analysis
AIMEMath96.9%58.916 Apr 2026Vals AI
OTIS Mock AIME 2024-2025Math96.1%65.3xhigh effortEpoch AI
MGSMMultilingual94.0%9 Jan 2026Vals AI
LiveBench MathematicsMath93.2%72.4high effort25 Jun 2026LiveBench
Graduate-Level Google-Proof Q&AKnowledge92.4%63.5David Rein et al.
GPQA DiamondKnowledge91.7%62.91 Sept 2026Vals AI
GPQA diamondKnowledge91.4%62.6xhigh effortEpoch AI
Artificial Analysis GPQA DiamondKnowledge90.3%61.1Artificial Analysis
τ²-bench TelecomAgentic89.7%63.2high effort · Sierra2 Mar 2026Sierra Research
MMMU ProMultimodal86.7%68.41 Sept 2026Vals AI
MMLU ProKnowledge86.2%56.41 Sept 2026Vals AI
ARC-AGI-1 (semi-private)Reasoning86.2%67.9xhigh effortARC Prize Foundation
LiveCodeBenchCoding85.4%61.91 Sept 2026Vals AI
τ²-Bench Tool-Agent-User EvaluationAgentic84.8%59.7Victor Barres et al.
LiveBench ReasoningReasoning83.2%71.6high effort25 Jun 2026LiveBench
MathVisionMultimodal83.0%Qwen
τ²-bench AirlineAgentic83.0%58.4high effort · Sierra2 Mar 2026Sierra Research
Artificial Analysis Long Context ReasoningReasoning82.7%65.5Artificial Analysis
CharXiv ReasoningMultimodal82.1%59.6CharXiv authors
τ²-bench RetailAgentic81.6%57.4high effort · Sierra30 Apr 2026Sierra Research
Software Engineering Benchmark VerifiedCoding80.0%61.7Carlos E. Jimenez et al.
LiveBench LanguageKnowledge79.8%66.8high effort25 Jun 2026LiveBench
Massive Multi-discipline Multimodal Understanding ProMultimodal79.5%56.7MMMU-Pro authors
LiveBench Data AnalysisReasoning78.2%64.5high effort25 Jun 2026LiveBench
LiveBench CodingCoding76.1%64.0high effort25 Jun 2026LiveBench
V*Multimodal75.9%31.9Z.AI
SWE-benchCoding75.8%58.31 Sept 2026Vals AI
Artificial Analysis IFBenchInstruction75.4%65.6Artificial Analysis
SWE-Bench verifiedCoding73.8%56.7high effortEpoch AI
SWE-bench VerifiedCoding72.3%55.5high effort · mini-SWE-agent1 Sept 2026SWE-bench team
FrontierMath-Tiers-1-3-v2-PrivateMath67.4%69.4xhigh effortEpoch AI
SWE-bench MultilingualMultilingual66.7%high effort · mini-SWE-agent20 Feb 2026SWE-bench team
SWE-bench MultilingualCoding66.7%high effort · mini-SWE-agent2 Sept 2026SWE-bench team
BrowseCompAgentic65.8%55.1OpenAI
Terminal-Bench 1.0Agentic63.8%65.612 Jan 2026Vals AI
EuroEval SwedishMultilingual63.5%81.6EuroEval
EuroEval PortugueseMultilingual62.5%80.4EuroEval
EuroEval FrenchMultilingual62.0%79.7EuroEval
LiveBench Instruction FollowingInstruction61.8%57.6high effort25 Jun 2026LiveBench
EuroEval ItalianMultilingual60.1%77.3EuroEval
EuroEval PolishMultilingual57.1%73.5EuroEval
EuroEval DutchMultilingual56.2%72.5EuroEval
SWE-bench ProCoding55.6%57.8Xiang Deng et al.
IOI v1Coding54.8%71.19 Aug 2026Vals AI
EuroEval SpanishMultilingual53.5%69.1EuroEval
Vibe Code Bench v1.1Coding53.5%64.4OpenHands21 Sept 2026Vals AI
ARC-AGI-2 (semi-private)Reasoning52.9%66.4xhigh effortARC Prize Foundation
Terminal-Bench 2.0Agentic51.7%59.64 Jun 2026Vals AI
LiveBench Agentic CodingAgentic50.3%64.0high effort25 Jun 2026LiveBench
EuroEval GermanMultilingual49.6%64.2EuroEval
Chess PuzzlesReasoning49.0%86.9xhigh effortEpoch AI
OSWorld-VerifiedAgentic47.3%37.7Tianbao Xie et al.
Gert Labs Composite Game BenchmarkAgentic46.5%53.8Gert Labs
Artificial Analysis Omniscience AccuracyKnowledge44.3%68.7Artificial Analysis
FrontierMath-2025-02-28-PrivateMath40.7%68.6xhigh effortEpoch AI
Furniture AssemblyReasoning38.3%66.8xhigh effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge37.7%65.4Artificial Analysis
SimpleQA VerifiedKnowledge37.1%55.9xhigh effortEpoch AI
JobBenchAgentic34.3%57.7Yuetai Li et al.
τ²-bench BankingAgentic32.2%22.0high effort · Sierra4 Aug 2026Sierra Research
FrontierMath-Tier-4-v2-PrivateMath31.7%63.2xhigh effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge30.4%60.5Artificial Analysis
EBR-benchReasoning23.0%65.1xhigh effortEpoch AI
Mystery Game PuzzlesReasoning23.0%58.4high effortEpoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath18.8%63.9xhigh effortEpoch AI
Critical Physics TasksReasoning11.6%62.7Artificial Analysis

44 benchmarks count, from 62 of 67 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsSierra Research, collected directlyMIT — results are in the licensed repositoryARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useSWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repository

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3