Granite 4.2 3B

Granite 4.2 3B is a reasoning model from IBM in the Granite 4.2 family. 17 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.32.8 ±5.4
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.128K
EloLMArena rating and rank.1297 (#238)

The index is a score out of 100. The ± range shows how much it can change.

3,072 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
35.5
CodingCode writing and repair.
33.8
ReasoningLogic problems and puzzles.
31.7
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
28.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
52.5
MathMath problems.
33.6

Results

17 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
American Invitational Mathematics Examination 2025Math78.3%38.8Mathematical Association of America
Instruction Following BenchmarkInstruction74.3%52.5Benchmark authors
LiveCodeBench v6Coding69.7%40.5LiveCodeBench maintainers
Massive Multitask Language Understanding ProfessionalKnowledge67.8%27.3Yubo Wang et al.
Harvard-MIT Mathematics Tournament February 2025Math66.7%28.3Qwen
Artificial Analysis GPQA DiamondKnowledge55.9%25.8Artificial Analysis
Graduate-Level Google-Proof Q&AKnowledge54.8%28.8David Rein et al.
Berkeley Function Calling Leaderboard v4Agentic52.4%36.9Arcee AI
τ³-Bench Tool-Agent-User EvaluationAgentic45.8%34.1Sierra Research
Artificial Analysis SciCodeCoding25.3%28.0Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning24.3%25.1Artificial Analysis
Scientific Code BenchmarkCoding24.1%33.0Benchmark authors
Artificial Analysis Omniscience AccuracyKnowledge9.2%25.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge9.1%33.7Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge6.6%31.7Artificial Analysis
GDPval-AA normalizedAgentic0.0%35.6Artificial Analysis
Critical Physics TasksReasoning0.0%38.4Artificial Analysis

17 benchmarks count, from 17 of 17 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

More from IBM

Granite 4.2 30B41.9Granite 4.2 8B37.6Granite-4.0-H-1B22.7Granite-4.0-H-350M20.6Granite-4.0-350M20.2