o1

o1 is a reasoning model from OpenAI. 22 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.45.4 ±3.9
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.16/s
Input / 1MUS dollars per 1M input tokens.$15
Output / 1MUS dollars per 1M output tokens.$60
ContextMaximum tokens in one request.200K
EloLMArena rating and rank.1366 (#177)

The index is a score out of 100. The ± range shows how much it can change.

27,807 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
43.7
CodingCode writing and repair.
38.5
ReasoningLogic problems and puzzles.
45.0
MultimodalTasks with images and text.
53.3
KnowledgeFacts and expert knowledge.
47.9
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
55.2
MathMath problems.
44.5

Results

22 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MATH level 5Math94.7%45.2high effortEpoch AI
Instruction-Following EvalInstruction92.2%49.9Jeffrey Zhou et al.
Massive Multitask Language UnderstandingKnowledge91.8%Dan Hendrycks et al.
MATH 500Math90.4%42.89 Jan 2026Vals AI
MGSMMultilingual89.3%9 Jan 2026Vals AI
MMLU ProKnowledge83.5%52.11 Sept 2026Vals AI
MMMU ProMultimodal77.4%53.31 Sept 2026Vals AI
GPQA diamondKnowledge76.8%49.1high effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge75.7%48.1David Rein et al.
Artificial Analysis GPQA DiamondKnowledge74.7%45.1Artificial Analysis
OTIS Mock AIME 2024-2025Math73.3%52.6high effortEpoch AI
GPQA DiamondKnowledge73.2%45.91 Sept 2026Vals AI
AIMEMath71.5%47.016 Apr 2026Vals AI
Artificial Analysis IFBenchInstruction70.3%60.4Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning65.0%53.2Artificial Analysis
τ²-Bench Tool-Agent-User EvaluationAgentic62.6%43.7Victor Barres et al.
LiveCodeBenchCoding50.3%30.01 Sept 2026Vals AI
SimpleQA VerifiedKnowledge41.1%59.6high effortEpoch AI
Artificial Analysis Coding IndexCoding39.7%47.0Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge34.5%56.5Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge15.2%41.4Artificial Analysis
Chess PuzzlesReasoning15.0%42.9high effortEpoch AI
FrontierMath-Tiers-1-3-v2-PrivateMath14.7%39.7high effortEpoch AI
FrontierMath-2025-02-28-PrivateMath9.3%39.3high effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge7.0%32.2Artificial Analysis
Critical Physics TasksReasoning0.3%39.0Artificial Analysis

22 benchmarks count, from 24 of 26 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AI

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3