Grok 4.3

Grok 4.3 is a reasoning model from xAI. 35 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.56.0 ±3.2
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.91/s
Input / 1MUS dollars per 1M input tokens.$1.25 batch $1
Output / 1MUS dollars per 1M output tokens.$2.5 batch $2 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1398 (#145)

The index is a score out of 100. The ± range shows how much it can change. Batch work costs less.

66,801 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
52.6
CodingCode writing and repair.
54.2
ReasoningLogic problems and puzzles.
50.6
MultimodalTasks with images and text.
60.5
KnowledgeFacts and expert knowledge.
60.9
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
63.8
MathMath problems.
60.6

Results

35 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic97.7%68.9Victor Barres et al.
GPQA DiamondKnowledge91.4%62.61 Sept 2026Vals AI
Graduate-Level Google-Proof Q&AKnowledge90.1%61.4David Rein et al.
Artificial Analysis GPQA DiamondKnowledge90.1%60.9Artificial Analysis
MMLU ProKnowledge85.8%55.81 Sept 2026Vals AI
LiveCodeBenchCoding84.5%61.11 Sept 2026Vals AI
LiveBench MathematicsMath84.3%60.625 Jun 2026LiveBench
MMMU ProMultimodal83.1%62.51 Sept 2026Vals AI
Instruction Following BenchmarkInstruction81.3%60.6Benchmark authors
Artificial Analysis IFBenchInstruction81.3%71.6Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal78.1%54.4MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal78.1%62.5Artificial Analysis
LiveBench LanguageKnowledge73.6%59.425 Jun 2026LiveBench
SWE-benchCoding71.4%54.81 Sept 2026Vals AI
LiveBench ReasoningReasoning70.8%54.325 Jun 2026LiveBench
LiveBench CodingCoding69.9%53.825 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning64.3%52.7Artificial Analysis
LiveBench Instruction FollowingInstruction62.8%59.225 Jun 2026LiveBench
LiveBench Data AnalysisReasoning55.8%33.425 Jun 2026LiveBench
Artificial Analysis SciCodeCoding48.3%59.9Artificial Analysis
Scientific Code BenchmarkCoding47.3%57.7Benchmark authors
Gert Labs Composite Game BenchmarkAgentic43.9%51.5Gert Labs
Terminal-Bench 2.0Agentic43.4%53.74 Jun 2026Vals AI
Artificial Analysis Coding IndexCoding42.3%48.7Artificial Analysis
Terminal-Bench 2.1Agentic41.9%48.821 Sept 2026Vals AI
SkillsBenchCoding40.6%57.5OpenHands11 Sept 2026Vals AI
Artificial Analysis Intelligence IndexKnowledge37.6%69.4Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge37.2%64.9Artificial Analysis
Humanity's Last ExamKnowledge35.0%58.4Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge34.6%56.7Artificial Analysis
GDPval-AA normalizedAgentic29.2%58.0Artificial Analysis
Vibe Code Bench v1.1Coding19.4%50.2OpenHands21 Sept 2026Vals AI
LiveBench Agentic CodingAgentic18.5%34.225 Jun 2026LiveBench
Artificial Analysis Agentic IndexAgentic17.2%51.2Artificial Analysis
APEX-Agents-AAAgentic17.0%54.5Artificial Analysis / Mercor
IOI v1Coding15.3%48.79 Aug 2026Vals AI
ResearchClawBenchAgentic12.4%InternScience
Critical Physics TasksReasoning8.0%55.2Artificial Analysis
Code MigrationCoding6.8%49.221 Sept 2026Vals AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI

35 benchmarks count, from 38 of 40 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not results

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from xAI

Grok 4.671.8Grok 4.770.1Grok 4.566.0Grok 4.20 (Reasoning)56.9Grok 451.6