GPT-6.1 Sol

GPT-6.1 Sol is a reasoning model from OpenAI in the GPT-6 family. 42 benchmarks count toward its index, in 8 categories.

  • Proprietary
  • Reasoning
  • Input: file · image · text
  • Released 2026-09
availableShows if the model has enough results for an index.
IndexOverall score out of 100.79.9±2.5Median 50.7
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.30/sMedian 62/s
First tokenSeconds to the first output token.5.8 sMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.$2Median $0.6
Output / 1MUS dollars per 1M output tokens.$10Median $2.5
ContextMaximum tokens in one request.1.05MMedian 262K
EloLMArena rating and rank.N/AMedian 1427

Details

From OpenRouter
API IDThe model ID on OpenRouter.
openai/gpt-6.1-sol
Max outputMaximum output tokens per request.
128K
Knowledge cutoffLast date of training data.
N/A
ToolsTool calls on OpenRouter.
Yes
JSON outputOutput follows a given JSON schema.
Yes
Cache read / 1MUS dollars per 1M cached tokens.
$0.1
WeightsLink to the published weights.
N/A
RetiresDate OpenRouter removes the model.
N/A

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.11 counted78.16 / 184
CodingCode writing and repair.5 counted76.08 / 165
ReasoningLogic problems and puzzles.11 counted82.65 / 195
MultimodalTasks with images and text.1 counted67.9—
KnowledgeFacts and expert knowledge.7 counted79.43 / 228
MultilingualTasks in many languages.1 counted95.0—
InstructionTasks with strict rules in the prompt.1 counted71.4—
MathMath problems.5 counted80.32 / 129

Results

42 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
OTIS Mock AIME 2024-2025Math100.0%1 / 185
Critical Physics TasksReasoning31.7%2 / 187
EuroEval DutchMultilingual74.0% ±0.42 / 167
GPQA diamond · Epoch AIKnowledge95.4% ±1.43 / 203
ARC-AGI-2 (semi-private)Reasoning94.2%2 / 122
Chess PuzzlesReasoning61.0% ±4.93 / 136
SimpleQA VerifiedKnowledge73.9% ±1.42 / 77
Mystery Game PuzzlesReasoning80.0% ±4.02 / 72
Artificial Analysis MMMU-ProMultimodal86.0%3 / 103
Agent Arena task outcomeAgentic15.52 / 48
Agent Arena steerabilityAgentic11.22 / 48
Artificial Analysis GDP.pdfAgentic31.0%1 / 23
LiveBench Data AnalysisReasoning82.2%3 / 62
Vibe Code Bench v1.1Coding88.9%7 / 107
ARC-AGI-1 (semi-private)Reasoning96.5%8 / 117
Code MigrationCoding65.1%5 / 73
Furniture AssemblyReasoning80.0% ±4.92 / 29
IOICoding96.9%3 / 41
LiveBench MathematicsMath96.5%5 / 62
LiveBench ReasoningReasoning91.6%6 / 62
LiveBench LanguageKnowledge88.6%6 / 62
ProofBench v1.1Math99.0%5 / 49
ExploitGymAgentic35.1%2 / 15
Terminal-Bench 4.0Agentic55.1%6 / 44
Terminal-Bench 4.0.0Agentic58.2% ±1.63 / 20
ARC-AGI-3 (semi-private)Reasoning52.7%4 / 25
GDPval-AA normalizedAgentic53.8%20 / 121
EBR-benchReasoning54.3% ±5.84 / 24
DeepSWEAgentic71.9%10 / 41
LiveBench CodingCoding80.7%16 / 62
Artificial Analysis SciCodeCoding54.2%28 / 103
HealthBench HardKnowledge36.2%4 / 12
LiveBench Instruction FollowingInstruction71.3%22 / 62
LiveBench Agentic CodingAgentic56.8%22 / 62
HealthBench ProfessionalKnowledge64.2%5 / 14
AutomationBenchAgentic36.1%13 / 22
ARC-AGI-2 (public eval)Reasoning97.1%4 / 113
ARC-AGI-1 (public eval)Reasoning98.5%4 / 107
Terminal-Bench ScienceScience52.9%2 / 38
SRE BenchOther50.8%2 / 32
ProgramBenchCoding3.0%5 / 61
BioMysteryBenchScience79.6%2 / 24
EMBFinance70.8%12 / 70
MedScribeHealthcare86.5%20 / 107
MysteryMechanismScience46.4%5 / 24
MedCodeHealthcare48.8%26 / 104
SAGEEducation46.5%35 / 90
Legal Research BenchLegal38.5%29 / 73
Tax Agent BenchFinance62.3%28 / 66
Finance Agent (v2)Finance52.0%32 / 74
Public Benefits Bench v1.1Public benefits59.3%32 / 47
CyberBench v1.1Other39.3%42 / 44
HealthBench raw scoreKnowledge56.7%—

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.aiOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

More from OpenAI

All models
GPT-6 Astra82.1GPT-5.5 Pro77.3GPT-6 Sol76.3GPT-5.4 Pro75.8GPT-5.6 Sol75.4