Granite 4.2 30B

Granite 4.2 30B is a reasoning model from IBM in the Granite 4.2 family. 20 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.41.9 ±5.1
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.128K
EloLMArena rating and rank.1363 (#181)

The index is a score out of 100. The ± range shows how much it can change.

3,255 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
43.3
CodingCode writing and repair.
43.9
ReasoningLogic problems and puzzles.
40.6
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
36.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
55.8
MathMath problems.
47.6

Results

20 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
American Invitational Mathematics Examination 2025Math89.2%47.0Mathematical Association of America
Harvard-MIT Mathematics Tournament February 2025Math89.2%48.3Qwen
Massive Multitask Language Understanding ProfessionalKnowledge77.6%42.8Yubo Wang et al.
Instruction Following BenchmarkInstruction77.2%55.8Benchmark authors
LiveCodeBench v6Coding75.8%46.0LiveCodeBench maintainers
Graduate-Level Google-Proof Q&AKnowledge66.4%39.6David Rein et al.
Artificial Analysis GPQA DiamondKnowledge64.4%34.6Artificial Analysis
τ³-Bench Tool-Agent-User EvaluationAgentic62.0%47.9Sierra Research
Berkeley Function Calling Leaderboard v4Agentic61.4%46.1Arcee AI
Software Engineering Benchmark VerifiedCoding57.0%43.2Carlos E. Jimenez et al.
Artificial Analysis Long Context ReasoningReasoning49.0%42.2Artificial Analysis
Scientific Code BenchmarkCoding38.8%48.6Benchmark authors
Artificial Analysis SciCodeCoding37.8%45.3Artificial Analysis
SWE-bench ProCoding33.3%36.2Xiang Deng et al.
Terminal-Bench 2.1 (provider run)Agentic29.2%41.3DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic29.2%41.3DeepSeek-AI
Artificial Analysis Intelligence IndexKnowledge14.8%40.9Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge11.2%36.7Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge10.1%26.4Artificial Analysis
GDPval-AA normalizedAgentic3.2%38.0Artificial Analysis
Critical Physics TasksReasoning0.3%39.0Artificial Analysis

20 benchmarks count, from 21 of 21 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

More from IBM

Granite 4.2 8B37.6Granite 4.2 3B32.8Granite-4.0-H-1B22.7Granite-4.0-H-350M20.6Granite-4.0-350M20.2