GPT-5.2-Codex

GPT-5.2-Codex is a reasoning model from OpenAI. 20 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.62.1 ±4.1
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.15/s
Input / 1MUS dollars per 1M input tokens.$1.75
Output / 1MUS dollars per 1M output tokens.$14
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
59.6
CodingCode writing and repair.
63.6
ReasoningLogic problems and puzzles.
62.0
MultimodalTasks with images and text.
60.3
KnowledgeFacts and expert knowledge.
61.2
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
66.4
MathMath problems.
66.5

Results

20 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic92.1%64.9Victor Barres et al.
Artificial Analysis GPQA DiamondKnowledge89.9%60.7Artificial Analysis
LiveBench MathematicsMath88.8%66.525 Jun 2026LiveBench
LiveCodeBenchCoding88.0%64.31 Sept 2026Vals AI
LiveBench CodingCoding83.6%76.425 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning82.3%65.2Artificial Analysis
LiveBench Data AnalysisReasoning78.2%64.625 Jun 2026LiveBench
LiveBench ReasoningReasoning77.7%63.925 Jun 2026LiveBench
Artificial Analysis IFBenchInstruction77.6%67.8Artificial Analysis
Artificial Analysis MMMU-ProMultimodal76.3%60.3Artificial Analysis
LiveBench LanguageKnowledge73.7%59.525 Jun 2026LiveBench
SWE-bench VerifiedCoding72.8%55.9mini-SWE-agent1 Sept 2026SWE-bench team
SWE-benchCoding72.4%55.61 Sept 2026Vals AI
LiveBench Instruction FollowingInstruction66.4%65.025 Jun 2026LiveBench
SWE-bench MultilingualCoding66.3%mini-SWE-agent2 Sept 2026SWE-bench team
Gert Labs Composite Game BenchmarkAgentic51.8%58.5Gert Labs
LiveBench Agentic CodingAgentic49.4%63.225 Jun 2026LiveBench
Artificial Analysis Omniscience AccuracyKnowledge41.1%64.7Artificial Analysis
Vibe Code Bench v1.1Coding37.9%57.9OpenHands21 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge35.7%63.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge28.5%58.0Artificial Analysis
JobBenchAgentic26.0%52.0Yuetai Li et al.
Critical Physics TasksReasoning8.7%56.6Artificial Analysis

20 benchmarks count, from 22 of 23 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be cited

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3