GPT-5.6 Sol

GPT-5.6 Sol is a reasoning model from OpenAI in the GPT-5.6 family. 59 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.76.7 ±2.1
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.38/s
Input / 1MUS dollars per 1M input tokens.$2 batch $1
Output / 1MUS dollars per 1M output tokens.$10 batch $5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1455 (#37)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

27,069 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
74.6
CodingCode writing and repair.
74.2
ReasoningLogic problems and puzzles.
81.2
MultimodalTasks with images and text.
68.1
KnowledgeFacts and expert knowledge.
76.5
MultilingualTasks in many languages.
89.0
InstructionTasks with strict rules in the prompt.
62.9
MathMath problems.
79.9

Results

59 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math100.0%67.4max effortEpoch AI
ARC-AGI-1 (semi-private)Reasoning96.5%72.8max effortARC Prize Foundation
SWE-benchCoding96.2%74.71 Sept 2026Vals AI
GPQA DiamondKnowledge95.2%66.11 Sept 2026Vals AI
Graduate-Level Google-Proof Q&AKnowledge94.6%65.6David Rein et al.
GPQA DiamondKnowledge94.6%65.6David Rein et al.
Artificial Analysis GPQA DiamondKnowledge94.1%65.0Artificial Analysis
GPQA diamondKnowledge93.5%64.6max effortEpoch AI
ARC-AGI-2 (semi-private)Reasoning92.5%86.4max effortARC Prize Foundation
BrowseCompAgentic92.2%77.1OpenAI
IOICoding91.2%85.421 Sept 2026Vals AI
FrontierMath-Tiers-1-3-v2-PrivateMath89.1%81.6max effortEpoch AI
MMLU ProKnowledge89.1%60.91 Sept 2026Vals AI
MMMU ProMultimodal88.8%71.91 Sept 2026Vals AI
Artificial Analysis Harvey LAB-AAAgentic87.2%68.8Artificial Analysis
VulcanBench v3Coding87.0%72.7VulcanBench contributors
IOI v1Coding86.7%89.29 Aug 2026Vals AI
VulcanBench Coding Intelligence Index v1Coding86.5%VulcanBench contributors
Terminal-Bench 2.1Agentic85.8%74.721 Sept 2026Vals AI
τ²-Bench Tool-Agent-User EvaluationAgentic85.1%59.9Victor Barres et al.
MMMU-Pro with PythonMultimodal84.6%OpenAI
CyberGymAgentic84.5%74.0Zhun Wang et al.
Artificial Analysis Long Context ReasoningReasoning84.0%66.4Artificial Analysis
Artificial Analysis MMMU-ProMultimodal83.4%69.0Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal83.0%62.4MMMU-Pro authors
ProofBench v1.1Math83.0%83.021 Sept 2026Vals AI
FrontierMath-Tier-4-v2-PrivateMath82.9%87.7max effortEpoch AI
LiveCodeBenchCoding82.6%59.41 Sept 2026Vals AI
LABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchKnowledge82.1%Jon M. Laurent et al.
Vibe Code Bench v1.1Coding80.5%75.7OpenHands21 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding77.4%73.5Artificial Analysis
EuroEval SwedishMultilingual74.9%95.0EuroEval
DeepSWEAgentic72.7%76.6Datacurve AI
Artificial Analysis IFBenchInstruction72.7%62.9Artificial Analysis
EuroEval FrenchMultilingual72.0%92.1EuroEval
EuroEval PortugueseMultilingual70.5%90.2EuroEval
EuroEval ItalianMultilingual70.1%89.8EuroEval
SimpleQA VerifiedKnowledge69.7%86.1max effortEpoch AI
EuroEval DutchMultilingual68.7%88.1EuroEval
cursorBench32Coding67.2%75.3Benchmark authors
EuroEval SpanishMultilingual66.3%85.0EuroEval
Terminal-Bench HardAgentic65.9%Artificial Analysis
SWE-bench ProCoding64.6%66.5Xiang Deng et al.
EuroEval GermanMultilingual63.6%81.7EuroEval
EuroEval PolishMultilingual62.8%80.7EuroEval
OSWorld 2.0Agentic62.6%85.7Mengqi Yuan et al.
FrontierCode 1.1 ExtendedCoding60.6%Cognition
HealthBench ProfessionalKnowledge60.5%Rebecca Soskin Hicks et al.
Artificial Analysis AutomationBenchAgentic60.1%72.1Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge59.4%87.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge58.9%95.0Artificial Analysis
ToolathlonAgentic58.0%72.1OpenAI
Mystery Game PuzzlesReasoning58.0%95.0max effortEpoch AI
Artificial Analysis SciCodeCoding57.1%72.1Artificial Analysis
Furniture AssemblyReasoning56.7%79.9max effortEpoch AI
Artificial Analysis ITBench-AAAgentic56.2%Artificial Analysis
Chess PuzzlesReasoning55.0%94.6max effortEpoch AI
HLE-VerifiedKnowledge54.5%Weiqi Zhai et al.
GDPval-AA normalizedAgentic54.4%77.3Artificial Analysis
SkillsBenchCoding54.1%68.9OpenHands11 Sept 2026Vals AI
Code MigrationCoding52.9%78.821 Sept 2026Vals AI
Artificial Analysis Agentic IndexAgentic50.5%78.4Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge49.5%78.2Artificial Analysis
Artificial Analysis AnalystAgentAgentic47.5%74.8Artificial Analysis
τ²-bench BankingAgentic46.9%32.5xhigh effort · Sierra4 Aug 2026Sierra Research
EBR-benchReasoning44.8%79.9max effortEpoch AI
Artificial Analysis Tau3-BankingAgentic44.3%72.8Artificial Analysis
OEIS Open LiteMath43.0%medium effortEpoch AI
Artificial Analysis EnterpriseOps-GymAgentic42.9%66.1Artificial Analysis
Bug Hunt BenchCoding42.0%Pawel Huryn
cursorBench40Coding41.7%Benchmark authors
Terminal-Bench 4.0.0Agentic37.3%85.7max effort · Codex21 Sept 2026Terminal-Bench
Terminal-Bench 3.0Agentic34.6%81.0Ryan Marten et al.
ExploitGymAgentic33.7%85.3Zhun Wang et al.
HealthBench HardKnowledge33.1%73.7Meta AI
Critical Physics TasksReasoning32.3%95.0Artificial Analysis
FrontierSWE v2Coding32.2%72.3Proximal
GeneBench-ProReasoning28.7%OpenAI
Terminal-Bench 4.0Agentic27.8%79.021 Sept 2026Vals AI
Artificial Analysis GDP.pdfAgentic27.2%78.8Artificial Analysis
Medical Long Context Reasoning (MLCR-AA)Reasoning26.1%60.5Wisedocs and Artificial Analysis
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic26.0%74.6NeoCognition
Vibe Code Bench 1-100Coding20.0%73.8OpenHands16 Sept 2026Vals AI
MirrorCodeCoding20.0%high effortEpoch AI
ARC-AGI-3 (semi-private)Reasoning7.8%max effortARC Prize Foundation
Agent Arena steerabilityAgentic6.674.8xhigh effort15 Sept 2026LMArena
Agent Arena command recoveryAgentic4.372.2xhigh effort15 Sept 2026LMArena
Agent Arena task outcomeAgentic3.771.6xhigh effort15 Sept 2026LMArena
ProgramBenchCoding1.5%21 Sept 2026Vals AI
FrontierMath-ErdosMath0.0%max effortEpoch AI

59 benchmarks count, from 74 of 90 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositorySierra Research, collected directlyMIT — results are in the licensed repositoryTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Muse Spark 1.375.0 · $4.25

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3GPT-5.5 Pro77.9