Kimi K3

Kimi K3 is a reasoning model from Moonshot AI. 38 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.72.0 ±4.6
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.37/s
Input / 1MUS dollars per 1M input tokens.$3
Output / 1MUS dollars per 1M output tokens.$15
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
72.6
CodingCode writing and repair.
68.4
ReasoningLogic problems and puzzles.
76.1
MultimodalTasks with images and text.
67.8
KnowledgeFacts and expert knowledge.
70.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

38 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MathVision with PythonMultimodal97.8%Moonshot AI / MathVision authors
DeepSearchQAAgentic95.0%73.6Meta AI
Artificial Analysis Harvey LAB-AAAgentic94.6%79.0Artificial Analysis
MathVisionMultimodal94.3%Qwen
Graduate-Level Google-Proof Q&AKnowledge93.5%64.6David Rein et al.
GPQA DiamondKnowledge93.5%64.6David Rein et al.
Artificial Analysis GPQA DiamondKnowledge93.5%64.4Artificial Analysis
CharXiv ReasoningMultimodal91.3%70.1CharXiv authors
BrowseCompAgentic91.2%76.3OpenAI
OmniDocBenchMultimodal91.1%Moonshot AI / OmniDocBench authors
Artificial Analysis Long Context ReasoningReasoning88.7%69.6Artificial Analysis
BabyVision with PythonMultimodal85.7%Moonshot AI
CharXiv Reasoning without toolsMultimodal84.8%CharXiv authors
MCP AtlasAgentic84.2%71.5OpenAI
MMMU-Pro with PythonMultimodal83.4%OpenAI
Massive Multi-discipline Multimodal Understanding ProMultimodal81.6%60.1MMMU-Pro authors
FrontierSWECoding81.2%Evan Chu et al.
Artificial Analysis MMMU-ProMultimodal80.5%65.5Artificial Analysis
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding77.8%75.3John Yang et al.
Artificial Analysis Coding IndexCoding76.2%72.7Artificial Analysis
VulcanBench v3Coding73.7%51.2VulcanBench contributors
DECK-Bench (Internal)Agentic73.5%Moonshot AI
Toolathlon-VerifiedAgentic73.2%71.5Moonshot AI
Kimi Code Bench v2Coding72.9%Moonshot AI
DeepSWEAgentic67.5%72.8Datacurve AI
OfficeQA ProMultimodal63.3%75.6OfficeQA Pro authors
cursorBench32Coding60.8%69.3Benchmark authors
Artificial Analysis SciCodeCoding59.5%75.4Artificial Analysis
PerceptionBench (Internal)Multimodal58.5%Moonshot AI
Artificial Analysis AutomationBenchAgentic58.3%69.8Artificial Analysis
OpenHarmony Bench v1.0Coding57.3%66.4OpenHarmony Bench authors
Humanity's Last ExamKnowledge56.0%76.2Center for AI Safety et al.
JobBenchAgentic52.9%70.5Yuetai Li et al.
GDPval-AA normalizedAgentic51.2%74.9Artificial Analysis
WorldVQA ForceAnswerMultimodal51.0%Moonshot AI / WorldVQA authors
Artificial Analysis Agentic IndexAgentic50.6%78.5Artificial Analysis
MLS-Bench LiteCoding48.3%MLS-Bench
Artificial Analysis ITBench-AAAgentic47.7%Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge47.6%72.8Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge46.9%75.4Artificial Analysis
Artificial Analysis Tau3-BankingAgentic46.0%75.1Artificial Analysis
Artificial Analysis EnterpriseOps-GymAgentic45.3%69.4Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge43.6%76.9Artificial Analysis
Humanity's Last Exam without toolsKnowledge43.5%65.6OpenAI
SWE-MarathonCoding42.0%Abundant AI and BenchFlow
APEX-Agents-AAAgentic41.3%73.7Artificial Analysis / Mercor
ZeroBench_main with PythonMultimodal41.0%Moonshot AI / ZeroBench authors
Artificial Analysis AnalystAgentAgentic38.8%68.7Artificial Analysis
Medical Long Context Reasoning (MLCR-AA)Reasoning38.3%71.2Wisedocs and Artificial Analysis
APEX-AgentsAgentic37.6%70.7Moonshot AI / APEX-Agents benchmark authors
PostTrain BenchCoding36.6%Moonshot AI
SpreadsheetBench 2Agentic34.8%Moonshot AI
AutomationBenchAgentic30.8%66.7Moonshot AI
FrontierSWE v2Coding25.9%68.8Proximal
Critical Physics TasksReasoning23.4%87.4Artificial Analysis
ZeroBenchMultimodal23.0%Meta AI
Artificial Analysis GDP.pdfAgentic22.0%73.2Artificial Analysis
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic18.0%70.1NeoCognition

38 benchmarks count, from 40 of 58 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from Moonshot AI

Kimi K2.660.9Kimi K2.7 Code59.5Kimi K2.5 (Reasoning)54.6Kimi K2.553.3Kimi K2.5 Thinking53.8