Moonshot AI
availableShows if the model has enough results for an index.Kimi K2 Instruct
Kimi K2 Instruct is a model from Moonshot AI. 9 benchmarks count toward its score, in 4 categories.
IndexOverall score out of 100.43.1 ±8.9
CoverageShare of the index weight with results.65%
SpeedOutput tokens per second.—
Input / 1MUS dollars per 1M input tokens.N/A
Output / 1MUS dollars per 1M output tokens.N/A
ContextMaximum tokens in one request.N/A
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change.
CapabilitiesScore per category, out of 100.
Out of 100Results
9 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| MATH 500 | Math | 94.2% | 47.1 | — | 9 Jan 2026 | Vals AI |
| MGSM | Multilingual | 90.9% | — | — | 9 Jan 2026 | Vals AI |
| MMLU Pro | Knowledge | 79.4% | 45.6 | — | 1 Sept 2026 | Vals AI |
| GPQA Diamond | Knowledge | 71.5% | 44.2 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 70.4% | 48.4 | — | 1 Sept 2026 | Vals AI |
| AIME | Math | 62.7% | 43.0 | — | 16 Apr 2026 | Vals AI |
| SWE-bench Verified | Coding | 48.6% | 36.5 | CodeSweep - SWE-agent | 1 Sept 2026 | SWE-bench team |
| SWE-bench Verified | Coding | 43.8% | 32.7 | mini-SWE-agent | 26 Feb 2026 | SWE-bench team |
| Terminal-Bench 1.0 | Agentic | 38.8% | 44.4 | — | 12 Jan 2026 | Vals AI |
| Terminal-Bench 1.0 | Agentic | 37.5% | 43.3 | — | 12 Jan 2026 | Vals AI |
| Terminal-Bench 2.0 | Agentic | 25.8% | 41.1 | — | 4 Jun 2026 | Vals AI |
| IOI v1 | Coding | 1.3% | 40.6 | — | 9 Aug 2026 | Vals AI |
9 benchmarks count, from 11 of 12 results. A grey row does not count. Too few models took that benchmark.