Inkling-Small

Inkling-Small is a hybrid model from Thinking Machines Lab in the Inkling family. 41 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.56.9 ±3.0
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.98/s
Input / 1MUS dollars per 1M input tokens.$0.45
Output / 1MUS dollars per 1M output tokens.$1.2
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1412 (#126)

The index is a score out of 100. The ± range shows how much it can change.

18,844 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
59.6
CodingCode writing and repair.
56.5
ReasoningLogic problems and puzzles.
55.1
MultimodalTasks with images and text.
54.7
KnowledgeFacts and expert knowledge.
56.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
61.7
MathMath problems.
56.6

Results

41 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIME 2026Math95.5%54.6Qwen
Harvard-MIT Mathematics Tournament February 2026Math90.2%57.7Qwen
OTIS Mock AIME 2024-2025Math90.0%61.9xhigh effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge89.5%60.9David Rein et al.
GPQA DiamondKnowledge89.5%60.9David Rein et al.
Artificial Analysis GPQA DiamondKnowledge89.5%60.3Artificial Analysis
GPQA diamondKnowledge88.5%60.0xhigh effortEpoch AI
LiveCodeBenchCoding85.9%62.41 Sept 2026Vals AI
MMLU ProKnowledge85.6%55.41 Sept 2026Vals AI
ARC-AGI-1 (semi-private)Reasoning84.0%66.9xhigh effortARC Prize Foundation
GPQA DiamondKnowledge83.6%55.41 Sept 2026Vals AI
Instruction Following BenchmarkInstruction82.2%61.7Benchmark authors
SWE-benchCoding82.2%63.51 Sept 2026Vals AI
CharXiv ReasoningMultimodal81.3%58.7CharXiv authors
Software Engineering Benchmark VerifiedCoding80.2%61.9Carlos E. Jimenez et al.
MCP AtlasAgentic79.6%68.1OpenAI
BrowseCompAgentic77.4%64.8OpenAI
CharXiv Reasoning without toolsMultimodal77.4%CharXiv authors
Artificial Analysis Long Context ReasoningReasoning75.7%60.6Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal74.0%47.8MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal74.0%57.5Artificial Analysis
SWE-bench ProCoding55.9%58.1Xiang Deng et al.
Terminal-Bench 2.1Agentic55.1%56.521 Sept 2026Vals AI
Toolathlon-VerifiedAgentic54.4%55.6Moonshot AI
Artificial Analysis Coding IndexCoding53.0%56.3Artificial Analysis
Artificial Analysis SciCodeCoding49.7%61.8Artificial Analysis
Scientific Code BenchmarkCoding48.7%59.2Benchmark authors
Humanity's Last ExamKnowledge47.8%69.3Center for AI Safety et al.
FrontierMath-Tiers-1-3-v2-PrivateMath46.3%57.5xhigh effortEpoch AI
ARC-AGI-2 (semi-private)Reasoning40.1%60.0xhigh effortARC Prize Foundation
SkillsBenchCoding33.6%51.5OpenHands11 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge33.3%60.7Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge33.2%54.9Artificial Analysis
Humanity's Last Exam without toolsKnowledge31.6%55.6OpenAI
Artificial Analysis Intelligence IndexKnowledge27.8%57.1Artificial Analysis
Artificial Analysis Agentic IndexAgentic24.9%57.5Artificial Analysis
SimpleQA VerifiedKnowledge19.1%39.2xhigh effortEpoch AI
Vibe Code Bench v1.1Coding19.1%50.1OpenHands21 Sept 2026Vals AI
Chess PuzzlesReasoning18.0%46.8xhigh effortEpoch AI
FrontierMath-Tier-4-v2-PrivateMath17.1%56.1xhigh effortEpoch AI
Code MigrationCoding13.7%53.721 Sept 2026Vals AI
IOICoding9.3%49.721 Sept 2026Vals AI
Critical Physics TasksReasoning8.3%55.8Artificial Analysis
Agent Arena command recoveryAgentic6.975.115 Sept 2026LMArena
ProofBench v1.1Math6.0%51.821 Sept 2026Vals AI
Mystery Game PuzzlesReasoning6.0%40.3xhigh effortEpoch AI
Terminal-Bench 4.0Agentic1.5%60.521 Sept 2026Vals AI
ProgramBenchCoding0.5%21 Sept 2026Vals AI
Agent Arena steerabilityAgentic-11.854.015 Sept 2026LMArena
Agent Arena task outcomeAgentic-20.444.415 Sept 2026LMArena

41 benchmarks count, from 48 of 50 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Thinking Machines Lab

Inkling56.7