OpenAI
availableShows if the model has enough results for an index.GPT-5.6 Sol
GPT-5.6 Sol is a reasoning model from OpenAI in the GPT-5.6 family. 59 benchmarks count toward its score, in 8 categories.
IndexOverall score out of 100.76.7 ±2.1
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.38/s
Input / 1MUS dollars per 1M input tokens.$2 batch $1
Output / 1MUS dollars per 1M output tokens.$10 batch $5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1455 (#37)
The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.
27,069 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
59 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 100.0% | 67.4 | max effort | — | Epoch AI |
| ARC-AGI-1 (semi-private) | Reasoning | 96.5% | 72.8 | max effort | — | ARC Prize Foundation |
| SWE-bench | Coding | 96.2% | 74.7 | — | 1 Sept 2026 | Vals AI |
| GPQA Diamond | Knowledge | 95.2% | 66.1 | — | 1 Sept 2026 | Vals AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 94.6% | 65.6 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 94.6% | 65.6 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 94.1% | 65.0 | — | — | Artificial Analysis |
| GPQA diamond | Knowledge | 93.5% | 64.6 | max effort | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 92.5% | 86.4 | max effort | — | ARC Prize Foundation |
| BrowseComp | Agentic | 92.2% | 77.1 | — | — | OpenAI |
| IOI | Coding | 91.2% | 85.4 | — | 21 Sept 2026 | Vals AI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 89.1% | 81.6 | max effort | — | Epoch AI |
| MMLU Pro | Knowledge | 89.1% | 60.9 | — | 1 Sept 2026 | Vals AI |
| MMMU Pro | Multimodal | 88.8% | 71.9 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Harvey LAB-AA | Agentic | 87.2% | 68.8 | — | — | Artificial Analysis |
| VulcanBench v3 | Coding | 87.0% | 72.7 | — | — | VulcanBench contributors |
| IOI v1 | Coding | 86.7% | 89.2 | — | 9 Aug 2026 | Vals AI |
| VulcanBench Coding Intelligence Index v1 | Coding | 86.5% | — | — | — | VulcanBench contributors |
| Terminal-Bench 2.1 | Agentic | 85.8% | 74.7 | — | 21 Sept 2026 | Vals AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 85.1% | 59.9 | — | — | Victor Barres et al. |
| MMMU-Pro with Python | Multimodal | 84.6% | — | — | — | OpenAI |
| CyberGym | Agentic | 84.5% | 74.0 | — | — | Zhun Wang et al. |
| Artificial Analysis Long Context Reasoning | Reasoning | 84.0% | 66.4 | — | — | Artificial Analysis |
| Artificial Analysis MMMU-Pro | Multimodal | 83.4% | 69.0 | — | — | Artificial Analysis |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 83.0% | 62.4 | — | — | MMMU-Pro authors |
| ProofBench v1.1 | Math | 83.0% | 83.0 | — | 21 Sept 2026 | Vals AI |
| FrontierMath-Tier-4-v2-Private | Math | 82.9% | 87.7 | max effort | — | Epoch AI |
| LiveCodeBench | Coding | 82.6% | 59.4 | — | 1 Sept 2026 | Vals AI |
| LABBench2: An Improved Benchmark for AI Systems Performing Biology Research | Knowledge | 82.1% | — | — | — | Jon M. Laurent et al. |
| Vibe Code Bench v1.1 | Coding | 80.5% | 75.7 | OpenHands | 21 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 77.4% | 73.5 | — | — | Artificial Analysis |
| EuroEval Swedish | Multilingual | 74.9% | 95.0 | — | — | EuroEval |
| DeepSWE | Agentic | 72.7% | 76.6 | — | — | Datacurve AI |
| Artificial Analysis IFBench | Instruction | 72.7% | 62.9 | — | — | Artificial Analysis |
| EuroEval French | Multilingual | 72.0% | 92.1 | — | — | EuroEval |
| EuroEval Portuguese | Multilingual | 70.5% | 90.2 | — | — | EuroEval |
| EuroEval Italian | Multilingual | 70.1% | 89.8 | — | — | EuroEval |
| SimpleQA Verified | Knowledge | 69.7% | 86.1 | max effort | — | Epoch AI |
| EuroEval Dutch | Multilingual | 68.7% | 88.1 | — | — | EuroEval |
| cursorBench32 | Coding | 67.2% | 75.3 | — | — | Benchmark authors |
| EuroEval Spanish | Multilingual | 66.3% | 85.0 | — | — | EuroEval |
| Terminal-Bench Hard | Agentic | 65.9% | — | — | — | Artificial Analysis |
| SWE-bench Pro | Coding | 64.6% | 66.5 | — | — | Xiang Deng et al. |
| EuroEval German | Multilingual | 63.6% | 81.7 | — | — | EuroEval |
| EuroEval Polish | Multilingual | 62.8% | 80.7 | — | — | EuroEval |
| OSWorld 2.0 | Agentic | 62.6% | 85.7 | — | — | Mengqi Yuan et al. |
| FrontierCode 1.1 Extended | Coding | 60.6% | — | — | — | Cognition |
| HealthBench Professional | Knowledge | 60.5% | — | — | — | Rebecca Soskin Hicks et al. |
| Artificial Analysis AutomationBench | Agentic | 60.1% | 72.1 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 59.4% | 87.3 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 58.9% | 95.0 | — | — | Artificial Analysis |
| Toolathlon | Agentic | 58.0% | 72.1 | — | — | OpenAI |
| Mystery Game Puzzles | Reasoning | 58.0% | 95.0 | max effort | — | Epoch AI |
| Artificial Analysis SciCode | Coding | 57.1% | 72.1 | — | — | Artificial Analysis |
| Furniture Assembly | Reasoning | 56.7% | 79.9 | max effort | — | Epoch AI |
| Artificial Analysis ITBench-AA | Agentic | 56.2% | — | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 55.0% | 94.6 | max effort | — | Epoch AI |
| HLE-Verified | Knowledge | 54.5% | — | — | — | Weiqi Zhai et al. |
| GDPval-AA normalized | Agentic | 54.4% | 77.3 | — | — | Artificial Analysis |
| SkillsBench | Coding | 54.1% | 68.9 | OpenHands | 11 Sept 2026 | Vals AI |
| Code Migration | Coding | 52.9% | 78.8 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Agentic Index | Agentic | 50.5% | 78.4 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 49.5% | 78.2 | — | — | Artificial Analysis |
| Artificial Analysis AnalystAgent | Agentic | 47.5% | 74.8 | — | — | Artificial Analysis |
| τ²-bench Banking | Agentic | 46.9% | 32.5 | xhigh effort · Sierra | 4 Aug 2026 | Sierra Research |
| EBR-bench | Reasoning | 44.8% | 79.9 | max effort | — | Epoch AI |
| Artificial Analysis Tau3-Banking | Agentic | 44.3% | 72.8 | — | — | Artificial Analysis |
| OEIS Open Lite | Math | 43.0% | — | medium effort | — | Epoch AI |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 42.9% | 66.1 | — | — | Artificial Analysis |
| Bug Hunt Bench | Coding | 42.0% | — | — | — | Pawel Huryn |
| cursorBench40 | Coding | 41.7% | — | — | — | Benchmark authors |
| Terminal-Bench 4.0.0 | Agentic | 37.3% | 85.7 | max effort · Codex | 21 Sept 2026 | Terminal-Bench |
| Terminal-Bench 3.0 | Agentic | 34.6% | 81.0 | — | — | Ryan Marten et al. |
| ExploitGym | Agentic | 33.7% | 85.3 | — | — | Zhun Wang et al. |
| HealthBench Hard | Knowledge | 33.1% | 73.7 | — | — | Meta AI |
| Critical Physics Tasks | Reasoning | 32.3% | 95.0 | — | — | Artificial Analysis |
| FrontierSWE v2 | Coding | 32.2% | 72.3 | — | — | Proximal |
| GeneBench-Pro | Reasoning | 28.7% | — | — | — | OpenAI |
| Terminal-Bench 4.0 | Agentic | 27.8% | 79.0 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis GDP.pdf | Agentic | 27.2% | 78.8 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 26.1% | 60.5 | — | — | Wisedocs and Artificial Analysis |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 26.0% | 74.6 | — | — | NeoCognition |
| Vibe Code Bench 1-100 | Coding | 20.0% | 73.8 | OpenHands | 16 Sept 2026 | Vals AI |
| MirrorCode | Coding | 20.0% | — | high effort | — | Epoch AI |
| ARC-AGI-3 (semi-private) | Reasoning | 7.8% | — | max effort | — | ARC Prize Foundation |
| Agent Arena steerability | Agentic | 6.6 | 74.8 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 4.3 | 72.2 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena task outcome | Agentic | 3.7 | 71.6 | xhigh effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 1.5% | — | — | 21 Sept 2026 | Vals AI |
| FrontierMath-Erdos | Math | 0.0% | — | max effort | — | Epoch AI |
59 benchmarks count, from 74 of 90 results. A grey row does not count. Too few models took that benchmark.