MiniCPM5-2B

MiniCPM5-2B is a reasoning model from OpenBMB in the MiniCPM5 family. 25 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.37.3 ±4.7
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.131K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
41.2
CodingCode writing and repair.
31.4
ReasoningLogic problems and puzzles.
44.0
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
35.2
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
40.7
MathMath problems.
44.5

Results

25 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MATH-500 Problem SetMath94.6%47.5Dan Hendrycks et al.
Instruction-Following EvalInstruction86.7%38.1Jeffrey Zhou et al.
American Invitational Mathematics Examination 2025Math86.5%45.0Mathematical Association of America
AIME 2026Math86.5%47.9Qwen
MMLU-ReduxKnowledge84.7%37.7Qwen
Massive Multitask Language Understanding ProfessionalKnowledge70.8%32.0Yubo Wang et al.
GPQA DiamondKnowledge70.2%43.1David Rein et al.
Artificial Analysis GPQA DiamondKnowledge70.2%40.5Artificial Analysis
LiveCodeBench v6Coding69.1%39.9LiveCodeBench maintainers
Berkeley Function Calling Leaderboard v4Agentic66.6%51.4Arcee AI
Instruction Following BenchmarkInstruction66.3%43.2Benchmark authors
Harvard-MIT Mathematics Tournament February 2026Math63.8%37.7Qwen
Artificial Analysis Long Context ReasoningReasoning59.0%49.1Artificial Analysis
Software Engineering Benchmark VerifiedCoding46.4%34.7Carlos E. Jimenez et al.
LongBench v2Reasoning43.7%LongBench v2 authors
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge40.8%31.1Xiaoxuan Du et al.
Scientific Code BenchmarkCoding26.3%35.3Benchmark authors
Artificial Analysis SciCodeCoding26.3%29.4Artificial Analysis
SWE-bench ProCoding14.4%17.9Xiang Deng et al.
Artificial Analysis Intelligence IndexKnowledge12.5%38.0Artificial Analysis
GDPval-AA normalizedAgentic9.7%43.0Artificial Analysis
Humanity's Last ExamKnowledge8.9%36.3Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge8.9%34.2Artificial Analysis
Terminal-Bench 2.1 (provider run)Agentic8.6%29.1DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic8.6%29.1DeepSeek-AI
Artificial Analysis Omniscience AccuracyKnowledge8.4%24.3Artificial Analysis
Critical Physics TasksReasoning0.3%39.0Artificial Analysis

25 benchmarks count, from 26 of 27 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

More from OpenBMB

MiniCPM5-1B14.2MiniCPM-o 2.6Unranked