DeepSeek V3 0324

DeepSeek V3 0324 is a standard model from DeepSeek in the DeepSeek family. 16 benchmarks count toward its index, in 7 categories.

  • Open weights
  • Standard
  • Input: text
  • Released 2025-03
availableShows if the model has enough results for an index.
IndexOverall score out of 100.35.8±5.0Median 50.7
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.—Median 62/s
First tokenSeconds to the first output token.N/AMedian 1.6 s
Input / 1MUS dollars per 1M input tokens.N/AMedian $0.6
Output / 1MUS dollars per 1M output tokens.N/AMedian $2.5
ContextMaximum tokens in one request.128KMedian 262K
EloLMArena rating and rank.1375±4 #165Median 1427

Capabilities

Score per category, out of 100.
Out of 100
AgenticMulti-step tasks with tools.2 counted31.7163 / 184
CodingCode writing and repair.4 counted39.8148 / 165
ReasoningLogic problems and puzzles.2 counted35.1158 / 195
MultimodalTasks with images and text.None countedN/A—
KnowledgeFacts and expert knowledge.4 counted37.2180 / 228
MultilingualTasks in many languages.1 counted37.6—
InstructionTasks with strict rules in the prompt.1 counted24.2—
MathMath problems.2 counted38.5105 / 129

Results

16 counted
BenchmarkThe test name.ResultThe score from the source.PlacePlace among results on this benchmark.
MATH 500Math88.6%30 / 59
AIMEMath52.2%60 / 95
MMLU ProKnowledge79.5%94 / 135
LiveCodeBenchCoding65.5%100 / 142
Critical Physics TasksReasoning0.0%136 / 187
Artificial Analysis IFBenchInstruction41.0%95 / 128
GDPval-AA normalizedAgentic0.0%91 / 121
Artificial Analysis SciCodeCoding39.0%79 / 103
Artificial Analysis GPQA DiamondKnowledge65.5%140 / 179
GPQA Diamond · Vals AIKnowledge61.6%109 / 135
MMLU-ProXMultilingual70.5%14 / 16
IOI v1Coding1.7%57 / 61
MGSMMultilingual91.7%26 / 74
TaxEval v2Finance71.1%83 / 143
CorpFin v2Finance54.7%92 / 131
LegalBenchLegal77.7%105 / 147
MedQAHealthcare82.0%68 / 95

Sources

BenchLM benchmark aggregationCC BY-NC 4.0 · Data from BenchLM.aiVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be cited

More from DeepSeek

All models
DeepSeek V4 Pro 081367.5DeepSeek V4.1 Flash67.0DeepSeek V4 Flash 073163.0DeepSeek V4 Pro 042358.5DeepSeek V3.2 (Thinking)49.1