Grok 4.7

Grok 4.7 is a reasoning model from xAI. 24 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.70.1 ±4.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.51/s
Input / 1MUS dollars per 1M input tokens.$1.6
Output / 1MUS dollars per 1M output tokens.$4.8
ContextMaximum tokens in one request.500K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
66.6
CodingCode writing and repair.
71.9
ReasoningLogic problems and puzzles.
67.9
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
72.9
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
78.8
MathMath problems.
67.9

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath95.7%75.8xhigh effort25 Jun 2026LiveBench
Vibe Code Bench v1.1Coding86.2%78.0OpenHands21 Sept 2026Vals AI
LiveBench ReasoningReasoning82.7%70.8xhigh effort25 Jun 2026LiveBench
LiveBench LanguageKnowledge80.1%67.2xhigh effort25 Jun 2026LiveBench
LiveBench CodingCoding77.2%65.8xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning76.9%62.8xhigh effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning76.7%61.3Artificial Analysis
LiveBench Instruction FollowingInstruction75.3%78.8xhigh effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic73.4%67.421 Sept 2026Vals AI
DeepSWEAgentic71.0%75.4Datacurve AI
Artificial Analysis AutomationBenchAgentic65.6%79.3Artificial Analysis
EEBench V1 core corpusCoding64.0%atopile
GDPval-AA normalizedAgentic59.8%81.5Artificial Analysis
IOICoding57.7%70.821 Sept 2026Vals AI
Artificial Analysis SciCodeCoding57.4%72.5Artificial Analysis
HealthBench ProfessionalKnowledge56.7%Rebecca Soskin Hicks et al.
LiveBench Agentic CodingAgentic54.0%67.5xhigh effort25 Jun 2026LiveBench
Artificial Analysis Omniscience AccuracyKnowledge47.4%72.5Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge46.5%80.5Artificial Analysis
cursorBench40Coding46.3%Benchmark authors
Code MigrationCoding44.8%73.621 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge43.1%71.3Artificial Analysis
Terminal-Bench 4.0.0Agentic37.6%85.9xhigh effort · Grok Build21 Sept 2026Terminal-Bench
FrontierSWE v2Coding29.5%70.8Proximal
ProofBench v1.1Math26.0%59.921 Sept 2026Vals AI
Artificial Analysis GDP.pdfAgentic20.0%71.1Artificial Analysis
Artificial Analysis Harvey LAB-AAAgentic19.6%5.0Artificial Analysis
Critical Physics TasksReasoning17.7%75.5Artificial Analysis

24 benchmarks count, from 25 of 28 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AITerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0

Same level, lower price

GPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87

More from xAI

Grok 4.671.8Grok 4.566.0Grok 4.356.0Grok 4.20 (Reasoning)56.9Grok 451.6