Mistral Large 4

Mistral Large 4 is a reasoning model from Mistral. 16 benchmarks count toward its index, in 6 categories.

  • Reasoning
  • Input: text
  • Released 2026-10
availableShows if the model has enough results for an index.
IndexOverall score out of 100.62.1±5.0Median 50.7
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.116/sMedian 62/s
First tokenSeconds to the first output token.18.7 sMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.$0.68Median $0.6
Output / 1MUS dollars per 1M output tokens.$2.09Median $2.5
ContextMaximum tokens in one request.1MMedian 262K
EloLMArena rating and rank.N/AMedian 1427

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.5 counted68.427 / 184
CodingCode writing and repair.5 counted65.630 / 165
ReasoningLogic problems and puzzles.2 counted61.148 / 195
MultimodalTasks with images and text.1 counted54.5—
KnowledgeFacts and expert knowledge.2 counted51.8106 / 228
MultilingualTasks in many languages.None countedN/A—
InstructionTasks with strict rules in the prompt.None countedN/A—
MathMath problems.1 counted50.8—

Results

16 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
Vibe Code Bench v1.1Coding78.4%25 / 107
GDPval-AA normalizedAgentic46.2%32 / 121
Artificial Analysis SciCodeCoding54.2%28 / 103
Critical Physics TasksReasoning10.6%51 / 187
Artificial Analysis MMMU-ProMultimodal76.4%37 / 103
Terminal-Bench 4.0Agentic22.7%19 / 44
Code MigrationCoding30.6%38 / 73
Artificial Analysis GDP.pdfAgentic18.6%14 / 23
IOICoding45.3%28 / 41
Vibe Code Bench 1-100Coding14.3%16 / 21
ProofBench v1.1Math10.0%44 / 49
Finance Agent (v2)Finance54.7%22 / 74
Tax Agent BenchFinance63.3%24 / 66
MedScribeHealthcare80.4%55 / 107
Legal Research BenchLegal31.7%38 / 73
MedCodeHealthcare40.7%59 / 104
EMBFinance56.2%43 / 70
BioMysteryBenchScience67.0%16 / 24
MysteryMechanismScience12.2%24 / 24
CybenchAgentic93.0%—

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.aiVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AI

Same level, lower price

Ling 3.1 Flash69.9 · Freedots3-note Preview65.0 · FreeMiMo-V2.6-Flash65.1 · $0.28

More from Mistral

All models
Mistral Medium 3.5 128B49.1Mistral Medium 3.541.9Magistral Medium38.0Mistral Small 436.5Magistral Small35.5