GPT-OSS 20B

GPT-OSS 20B is a non-reasoning model from OpenAI in the GPT-OSS family. 22 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.40.5 ±4.4
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.222/s
Input / 1MUS dollars per 1M input tokens.$0.03
Output / 1MUS dollars per 1M output tokens.$0.13
ContextMaximum tokens in one request.131K
EloLMArena rating and rank.1287 (#246)

The index is a score out of 100. The ± range shows how much it can change.

10,393 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
39.4
CodingCode writing and repair.
46.2
ReasoningLogic problems and puzzles.
32.4
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
34.6
MultilingualTasks in many languages.
61.9
InstructionTasks with strict rules in the prompt.
55.2
MathMath problems.
47.0

Results

22 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MATH 500Math94.2%47.19 Jan 2026Vals AI
MGSMMultilingual89.0%9 Jan 2026Vals AI
AIMEMath86.0%53.816 Apr 2026Vals AI
LiveCodeBenchCoding80.4%57.41 Sept 2026Vals AI
MMLU ProKnowledge71.6%33.31 Sept 2026Vals AI
React Native EvalsCoding71.0%47.1Callstack
GPQA DiamondKnowledge68.9%41.91 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge68.8%39.1Artificial Analysis
Artificial Analysis IFBenchInstruction65.1%55.2Artificial Analysis
τ²-Bench Tool-Agent-User EvaluationAgentic60.2%42.0Victor Barres et al.
EuroEval SwedishMultilingual55.5%71.6medium effortEuroEval
EuroEval FrenchMultilingual52.9%68.3medium effortEuroEval
EuroEval SwedishMultilingual52.0%67.3low effortEuroEval
EuroEval FrenchMultilingual51.7%66.8low effortEuroEval
EuroEval ItalianMultilingual51.0%66.0medium effortEuroEval
OTIS Mock AIME 2024-2025Math50.8%40.0high effortEpoch AI
EuroEval PortugueseMultilingual50.6%65.5medium effortEuroEval
EuroEval PolishMultilingual50.2%64.9medium effortEuroEval
EuroEval DutchMultilingual49.9%64.6low effortEuroEval
EuroEval PortugueseMultilingual49.7%64.3low effortEuroEval
EuroEval PolishMultilingual48.6%62.9low effortEuroEval
EuroEval SwedishMultilingual48.4%62.7high effortEuroEval
EuroEval ItalianMultilingual47.8%61.9low effortEuroEval
GPQA diamondKnowledge46.0%20.7high effortEpoch AI
EuroEval SpanishMultilingual45.8%59.5medium effortEuroEval
EuroEval GermanMultilingual45.8%59.4medium effortEuroEval
EuroEval FrenchMultilingual44.8%58.2high effortEuroEval
EuroEval SpanishMultilingual44.2%57.5low effortEuroEval
EuroEval PolishMultilingual43.5%56.6high effortEuroEval
EuroEval ItalianMultilingual43.3%56.4high effortEuroEval
EuroEval GermanMultilingual42.5%55.3low effortEuroEval
EuroEval PortugueseMultilingual40.7%53.1high effortEuroEval
EuroEval DutchMultilingual40.5%52.9high effortEuroEval
Artificial Analysis SciCodeCoding38.9%46.9Artificial Analysis
EuroEval GermanMultilingual38.3%50.1high effortEuroEval
Artificial Analysis Long Context ReasoningReasoning34.7%32.3Artificial Analysis
EuroEval SpanishMultilingual32.9%43.5high effortEuroEval
Artificial Analysis Coding IndexCoding20.7%33.6Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge16.0%33.7Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge11.0%36.5Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge9.0%33.6Artificial Analysis
Artificial Analysis Agentic IndexAgentic1.4%38.3Artificial Analysis
Critical Physics TasksReasoning1.4%41.3Artificial Analysis
APEX-Agents-AAAgentic0.7%41.6Artificial Analysis / Mercor
GDPval-AA normalizedAgentic0.0%35.6Artificial Analysis
Chess PuzzlesReasoning0.0%23.5high effortEpoch AI

22 benchmarks count, from 45 of 46 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryEpoch AI, collected directlyCC BY — free to use and redistribute with attribution

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3