Thinking Machines Lab
availableShows if the model has enough results for an index.Inkling
Inkling is a hybrid model from Thinking Machines Lab. 50 benchmarks count toward its score, in 7 categories.
IndexOverall score out of 100.56.7 ±2.8
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.119/s
Input / 1MUS dollars per 1M input tokens.$1
Output / 1MUS dollars per 1M output tokens.$4.05
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1440 (#67)
The index is a score out of 100. The ± range shows how much it can change. A free tier is available.
25,922 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category, out of 100.
Out of 100Results
50 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| AIME 2026 | Math | 97.1% | 55.8 | — | — | Qwen |
| OTIS Mock AIME 2024-2025 | Math | 88.9% | 61.2 | xhigh effort | — | Epoch AI |
| LiveBench Mathematics | Math | 88.4% | 66.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| GPQA diamond | Knowledge | 88.3% | 59.7 | xhigh effort | — | Epoch AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 87.9% | 59.4 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 87.9% | 59.4 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 87.2% | 57.9 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 87.1% | 58.7 | — | 1 Sept 2026 | Vals AI |
| MMLU Pro | Knowledge | 86.3% | 56.5 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 85.5% | 62.1 | — | 1 Sept 2026 | Vals AI |
| CharXiv Reasoning | Multimodal | 82.0% | 59.5 | — | — | CharXiv authors |
| Instruction Following Benchmark | Instruction | 79.8% | 58.9 | — | — | Benchmark authors |
| ARC-AGI-1 (semi-private) | Reasoning | 79.5% | 64.8 | — | — | ARC Prize Foundation |
| MMMU Pro | Multimodal | 78.5% | 55.1 | — | 1 Sept 2026 | Vals AI |
| LiveBench Reasoning | Reasoning | 78.3% | 64.8 | xhigh effort | 25 Jun 2026 | LiveBench |
| CharXiv Reasoning without tools | Multimodal | 78.1% | — | — | — | CharXiv authors |
| Software Engineering Benchmark Verified | Coding | 77.6% | 59.8 | — | — | Carlos E. Jimenez et al. |
| SWE-bench | Coding | 77.6% | 59.8 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 77.3% | 61.7 | — | — | Artificial Analysis |
| BrowseComp | Agentic | 77.1% | 64.6 | — | — | OpenAI |
| MCP Atlas | Agentic | 74.1% | 64.0 | — | — | OpenAI |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 73.5% | 47.0 | — | — | MMMU-Pro authors |
| Artificial Analysis MMMU-Pro | Multimodal | 73.5% | 56.9 | — | — | Artificial Analysis |
| LiveBench Language | Knowledge | 73.5% | 59.2 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Data Analysis | Reasoning | 72.8% | 57.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Coding | Coding | 71.0% | 55.6 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveBench Instruction Following | Instruction | 70.1% | 70.7 | xhigh effort | 25 Jun 2026 | LiveBench |
| SWE-bench Pro | Coding | 54.3% | 56.5 | — | — | Xiang Deng et al. |
| Artificial Analysis Coding Index | Coding | 52.1% | 55.6 | — | — | Artificial Analysis |
| LiveBench Agentic Coding | Agentic | 49.4% | 63.2 | xhigh effort | 25 Jun 2026 | LiveBench |
| Terminal-Bench 2.1 | Agentic | 47.6% | 52.1 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis SciCode | Coding | 47.0% | 58.1 | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 46.0% | 67.7 | — | — | Center for AI Safety et al. |
| Artificial Analysis Omniscience Accuracy | Knowledge | 41.6% | 65.3 | — | — | Artificial Analysis |
| SimpleQA Verified | Knowledge | 40.3% | 58.9 | xhigh effort | — | Epoch AI |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 38.0% | 59.4 | — | — | Artificial Analysis |
| ARC-AGI-2 (semi-private) | Reasoning | 36.5% | 58.1 | — | — | ARC Prize Foundation |
| FrontierMath-Tiers-1-3-v2-Private | Math | 33.3% | 50.2 | xhigh effort | — | Epoch AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 31.9% | 59.1 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 30.0% | 54.2 | — | — | OpenAI |
| Artificial Analysis Tau3-Banking | Agentic | 29.1% | 52.3 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 28.2% | 57.2 | — | — | Artificial Analysis |
| SkillsBench | Coding | 26.1% | 45.1 | OpenHands | 11 Sept 2026 | Vals AI |
| τ²-bench Banking | Agentic | 25.0% | 16.8 | max effort · Sierra | 4 Aug 2026 | Sierra Research |
| Artificial Analysis Intelligence Index | Knowledge | 25.0% | 53.6 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 24.3% | 57.0 | — | — | Artificial Analysis |
| Artificial Analysis AnalystAgent | Agentic | 23.8% | 58.2 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 21.0% | 50.7 | xhigh effort | — | Epoch AI |
| Vibe Code Bench v1.1 | Coding | 19.2% | 50.2 | OpenHands | 21 Sept 2026 | Vals AI |
| IOI | Coding | 14.9% | 52.1 | — | 21 Sept 2026 | Vals AI |
| Code Migration | Coding | 11.8% | 52.4 | — | 21 Sept 2026 | Vals AI |
| Vibe Code Bench 1-100 | Coding | 7.3% | 59.8 | OpenHands | 16 Sept 2026 | Vals AI |
| Critical Physics Tasks | Reasoning | 5.4% | 49.7 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 4.9% | 50.3 | xhigh effort | — | Epoch AI |
| FrontierSWE v2 | Coding | 4.1% | 56.8 | — | — | Proximal |
| Agent Arena command recovery | Agentic | 2.0 | 69.6 | — | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
| ProofBench v1.1 | Math | 0.0% | 49.4 | — | 21 Sept 2026 | Vals AI |
| Terminal-Bench 4.0 | Agentic | 0.0% | 59.5 | — | 21 Sept 2026 | Vals AI |
| Agent Arena steerability | Agentic | -10.5 | 55.5 | — | 15 Sept 2026 | LMArena |
| Agent Arena task outcome | Agentic | -19.3 | 45.6 | — | 15 Sept 2026 | LMArena |
50 benchmarks count, from 59 of 61 results. A grey row does not count. Too few models took that benchmark.