GPT-6 Sol

GPT-6 Sol is a reasoning model from OpenAI in the GPT-6 family. 16 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.76.9 ±6.0
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.$2
Output / 1MUS dollars per 1M output tokens.$10
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
78.0
CodingCode writing and repair.
72.8
ReasoningLogic problems and puzzles.
80.6
MultimodalTasks with images and text.
68.9
KnowledgeFacts and expert knowledge.
77.4
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

16 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Artificial Analysis Long Context ReasoningReasoning83.7%66.1Artificial Analysis
Artificial Analysis MMMU-ProMultimodal83.3%68.9Artificial Analysis
DeepSWEAgentic68.8%73.8Datacurve AI
Artificial Analysis AutomationBenchAgentic61.6%74.1Artificial Analysis
HealthBench ProfessionalKnowledge60.8%Rebecca Soskin Hicks et al.
OSWorld 2.0Agentic60.5%84.7Mengqi Yuan et al.
HealthBench Professional raw scoreKnowledge59.5%Anthropic
Artificial Analysis SciCodeCoding57.6%72.8Artificial Analysis
Agents' Last ExamAgentic56.4%94.3DeepSeek-AI
Artificial Analysis Omniscience AccuracyKnowledge54.5%81.3Artificial Analysis
HealthBench length-adjusted scoreKnowledge53.2%Anthropic
GDPval-AA normalizedAgentic49.3%73.4Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge47.9%76.5Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge47.5%81.8Artificial Analysis
HealthBench raw scoreKnowledge47.1%Anthropic
AutomationBenchAgentic33.2%70.8Moonshot AI
Critical Physics TasksReasoning30.9%95.0Artificial Analysis
HealthBench HardKnowledge30.1%69.8Meta AI
Artificial Analysis GDP.pdfAgentic24.8%76.2Artificial Analysis
ExploitGymAgentic22.1%77.0Zhun Wang et al.

16 benchmarks count, from 16 of 20 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Muse Spark 1.375.0 · $4.25

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-5.572.3GPT-5.5 Pro77.9