Step 5 Preview
Released Sep 20, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 64K
- Speed
- 92 tok/s · 3.05s to first answer token
- Price ($/MTok in / out)
- $1 / $2.70
- Cost per task
- $0.72
- Open weights
- No
Benchmarks1 of 10 reported
GPQA Diamond
MMLU-Proretired
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Sustains long agentic runs without hand-holding — StepFun's own long-horizon kernel-optimisation experiment had it iterating to a marginally faster kernel than Claude Opus 5 reached on the same task
- Streams quickly for a 600B reasoning model — a 1,000-token answer lands in about ten seconds where slower peers take half a minute, which matters once outputs run long
- Takes up to 60 images or a short video in a single request, so UI and document-heavy agent tasks stay in one thread
- Built with finance in mind: corporate valuation, live-search and financial deep-research tasks are where StepFun claims parity with the closed flagships
Weaknesses
- Very verbose in its reasoning: Artificial Analysis' index run drew nearly twice the median output tokens, so the cheap per-token rate buys less than it looks
- StepFun's comparison table runs Step 5 at high effort against rivals at max and mixes HLE tool settings, so its own parity claims are not apples-to-apples
- A preview with no model card, no weights yet and no stated licence — the open-weights promise for October 15 currently has nothing behind it but an empty Hugging Face repo
600B-total / 27B-active sparse MoE with 92 layers, StepFun's first Step 5 model, announced on X at 03:15 UTC on 2026-09-20 with API access the same day (Artificial Analysis dates it 2026-09-18). GPQA Diamond 93.5 is from StepFun's launch table, published only as an image and transcribed by CellCog and a linux.do relay, at high effort. That table also lists HLE 46.5%, but StepFun's footnote says its HLE row is a with-tools row mixing text-only-subset and full-dataset runs, so it is not recorded; it further lists Terminal-Bench 4.0 33.3% (not the 2.1 this site tracks), DeepSWE v1.1 67.7, SWE-Marathon v1.1 72.7, BrowseComp 88.7, FrontierFinance 66.4, DRACO 83.3, StepCodeBench 49.0, ProgramBench 80.5, ALE-CLI 29.5, MMMU-Pro 76.0 and a GDPval-AA figure of 1571 whose version label could not be confirmed; AA rates it 1566 on GDPval-AA v2.1, a rescaled successor to the v2 this site's gdpvalAA key records, so that cell stays null. AA Intelligence Index 44 (level with Grok 4.6 and Kimi K3 Max); costPerTask $0.72 and speed (92.1 tok/s, 3.05 s to first answer token) are AA's measurements of the single reasoning variant it lists, which carries no effort label even though the API exposes low/medium/high — hence effort null. Pricing is $1.00 input on cache miss / $0.05 on cache hit / $2.70 output including reasoning tokens; max output 64K per StepFun's platform docs. Weights are promised for 2026-10-15 in BF16 (~1.2 TB) with no licence named, so openWeights stays false until they ship. No arena.ai listing, no ARC Prize entry surfaced, knowledge cutoff undisclosed. predecessorId is null: StepFun frames it as a new flagship, not a replacement for Step 3.7 Flash.
For developers
API model strings
- stepfun
step-5-preview
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor