Step 5 Preview

Step 5 Preview is a reasoning model from StepFun in the Step 5 family. 24 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.72.5 ±5.3
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.83/s
Input / 1MUS dollars per 1M input tokens.$1
Output / 1MUS dollars per 1M output tokens.$2.7
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
74.6
CodingCode writing and repair.
73.9
ReasoningLogic problems and puzzles.
75.8
MultimodalTasks with images and text.
61.4
KnowledgeFacts and expert knowledge.
70.0
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
GPQA DiamondKnowledge93.5%64.6David Rein et al.
BrowseCompAgentic88.7%74.2OpenAI
Artificial Analysis Long Context ReasoningReasoning88.3%69.3Artificial Analysis
MCP AtlasAgentic85.6%72.6OpenAI
Terminal-Bench 2.1 (provider run)Agentic85.0%74.2DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic85.0%74.2DeepSeek-AI
CyberGymAgentic84.7%74.1Zhun Wang et al.
Data Research and Analysis with Complex OperationsAgentic83.3%Anthropic
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding80.5%77.1John Yang et al.
Artificial Analysis MMMU-ProMultimodal76.4%60.5Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal76.0%51.0MMMU-Pro authors
Toolathlon-VerifiedAgentic74.1%72.3Moonshot AI
SWE-MarathonCoding72.7%Abundant AI and BenchFlow
DeepSWEAgentic67.7%73.0Datacurve AI
OfficeQA ProMultimodal60.3%72.6OfficeQA Pro authors
JobBenchAgentic59.0%74.7Yuetai Li et al.
Scientific Code BenchmarkCoding58.9%70.0Benchmark authors
Artificial Analysis SciCodeCoding58.9%74.6Artificial Analysis
GDPval-AA normalizedAgentic53.3%76.5Artificial Analysis
Humanity's Last ExamKnowledge46.5%68.2Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge46.5%75.0Artificial Analysis
AutomationBenchAgentic44.0%89.2Moonshot AI
Artificial Analysis Intelligence IndexKnowledge43.7%77.1Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge41.5%65.2Artificial Analysis
MLS-Bench LiteCoding40.5%MLS-Bench
APEX-AgentsAgentic37.8%70.8Moonshot AI / APEX-Agents benchmark authors
Agents' Last ExamAgentic29.5%69.0DeepSeek-AI
Critical Physics TasksReasoning20.9%82.2Artificial Analysis

24 benchmarks count, from 25 of 28 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from StepFun

Step 3.7 Flash56.4Step 3.5 FlashUnrankedStep-AudioUnrankedStep-Audio-Chat 130BUnranked