DeepSeek
availableShows if the model has enough results for an index.DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a reasoning model from DeepSeek. 29 benchmarks count toward its score, in 6 categories.
IndexOverall score out of 100.70.6 ±4.0
CoverageShare of the index weight with results.90%
SpeedOutput tokens per second.103/s
Input / 1MUS dollars per 1M input tokens.$0.15
Output / 1MUS dollars per 1M output tokens.$0.6
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change.
CapabilitiesScore per category, out of 100.
Out of 100Results
29 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Graduate-Level Google-Proof Q&A | Knowledge | 90.9% | 62.2 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 90.9% | 62.2 | — | — | David Rein et al. |
| Terminal-Bench 2.1 (provider run) | Agentic | 90.6% | 77.5 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 90.6% | 77.5 | — | — | DeepSeek-AI |
| BabyVision with Python | Multimodal | 89.6% | — | — | — | Moonshot AI |
| CyberGym | Agentic | 88.1% | 76.4 | — | — | Zhun Wang et al. |
| Vibe Code Bench v1.1 | Coding | 84.7% | 77.4 | OpenHands | 21 Sept 2026 | Vals AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 84.0% | 66.4 | — | — | Artificial Analysis |
| Chartography with image and code tools | Multimodal | 78.9% | — | — | — | Surge AI and Anthropic |
| Artificial Analysis MMMU-Pro | Multimodal | 77.0% | 61.2 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 74.5% | 68.0 | — | 21 Sept 2026 | Vals AI |
| DeepSWE | Agentic | 74.2% | 77.7 | — | — | Datacurve AI |
| SkillsBench | Coding | 69.8% | 82.1 | OpenHands | 11 Sept 2026 | Vals AI |
| Artificial Analysis AutomationBench | Agentic | 68.9% | 83.7 | — | — | Artificial Analysis |
| Apex | Math | 65.6% | — | — | — | DeepSeek-AI |
| NL2Repo | Coding | 65.4% | 76.5 | — | — | MiniMax |
| Humanity's Last Exam with tools | Agentic | 63.9% | 76.4 | — | — | DeepSeek-AI |
| GDPval-AA normalized | Agentic | 55.0% | 77.8 | — | — | Artificial Analysis |
| AutomationBench | Agentic | 54.8% | 95.0 | — | — | Moonshot AI |
| ProofBench v1.1 | Math | 54.0% | 71.3 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis SciCode | Coding | 51.9% | 64.9 | — | — | Artificial Analysis |
| ZeroBench_main with Python | Multimodal | 49.0% | — | — | — | Moonshot AI / ZeroBench authors |
| Artificial Analysis Omniscience Accuracy | Knowledge | 46.4% | 71.3 | — | — | Artificial Analysis |
| Code Migration | Coding | 45.6% | 74.1 | — | 21 Sept 2026 | Vals AI |
| IOI | Coding | 40.3% | 63.2 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Intelligence Index | Knowledge | 39.5% | 71.7 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 39.2% | 67.1 | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 36.8% | 60.0 | — | — | Center for AI Safety et al. |
| Agents' Last Exam | Agentic | 31.8% | 71.1 | — | — | DeepSeek-AI |
| terminalBench3 | Agentic | 30.0% | — | — | — | Benchmark authors |
| terminalBench3 | Agentic | 30.0% | — | — | — | Benchmark authors |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 22.8% | 57.6 | — | — | Wisedocs and Artificial Analysis |
| ProgramBench: Can Language Models Rebuild Programs From Scratch? | Coding | 20.3% | 35.4 | — | — | John Yang et al. |
| Vibe Code Bench 1-100 | Coding | 16.4% | 69.8 | OpenHands | 16 Sept 2026 | Vals AI |
| ExploitGym | Agentic | 15.3% | 72.2 | — | — | Zhun Wang et al. |
| Critical Physics Tasks | Reasoning | 14.3% | 68.4 | — | — | Artificial Analysis |
| Agent Arena task outcome | Agentic | 13.4 | 82.4 | max effort | 15 Sept 2026 | LMArena |
| Terminal-Bench 4.0 | Agentic | 11.6% | 67.7 | — | 21 Sept 2026 | Vals AI |
| Agent Arena command recovery | Agentic | 7.6 | 76.0 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | -1.6 | 65.5 | max effort | 15 Sept 2026 | LMArena |
29 benchmarks count, from 34 of 40 results. A grey row does not count. Too few models took that benchmark.