Wait Which Model?
← Back to directory
StepFun

Step 5 Preview

Released Sep 20, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
China
Modality
Multimodal
Context window
1M
Max output
64K
Speed
92 tok/s · 3.05s to first answer token
Price ($/MTok in / out)
$1 / $2.70
Cost per task
$0.72
Open weights
No
Benchmarks1 of 10 reported

GPQA Diamond

93.5%

MMLU-Proretired

SWE-bench Verified

SWE-bench Pro

Terminal-Bench 2.1

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Sustains long agentic runs without hand-holding — StepFun's own long-horizon kernel-optimisation experiment had it iterating to a marginally faster kernel than Claude Opus 5 reached on the same task
  • Streams quickly for a 600B reasoning model — a 1,000-token answer lands in about ten seconds where slower peers take half a minute, which matters once outputs run long
  • Takes up to 60 images or a short video in a single request, so UI and document-heavy agent tasks stay in one thread
  • Built with finance in mind: corporate valuation, live-search and financial deep-research tasks are where StepFun claims parity with the closed flagships

Weaknesses

  • Very verbose in its reasoning: Artificial Analysis' index run drew nearly twice the median output tokens, so the cheap per-token rate buys less than it looks
  • StepFun's comparison table runs Step 5 at high effort against rivals at max and mixes HLE tool settings, so its own parity claims are not apples-to-apples
  • A preview with no model card, no weights yet and no stated licence — the open-weights promise for October 15 currently has nothing behind it but an empty Hugging Face repo

600B-total / 27B-active sparse MoE with 92 layers, StepFun's first Step 5 model, announced on X at 03:15 UTC on 2026-09-20 with API access the same day (Artificial Analysis dates it 2026-09-18). GPQA Diamond 93.5 is from StepFun's launch table, published only as an image and transcribed by CellCog and a linux.do relay, at high effort. That table also lists HLE 46.5%, but StepFun's footnote says its HLE row is a with-tools row mixing text-only-subset and full-dataset runs, so it is not recorded; it further lists Terminal-Bench 4.0 33.3% (not the 2.1 this site tracks), DeepSWE v1.1 67.7, SWE-Marathon v1.1 72.7, BrowseComp 88.7, FrontierFinance 66.4, DRACO 83.3, StepCodeBench 49.0, ProgramBench 80.5, ALE-CLI 29.5, MMMU-Pro 76.0 and a GDPval-AA figure of 1571 whose version label could not be confirmed; AA rates it 1566 on GDPval-AA v2.1, a rescaled successor to the v2 this site's gdpvalAA key records, so that cell stays null. AA Intelligence Index 44 (level with Grok 4.6 and Kimi K3 Max); costPerTask $0.72 and speed (92.1 tok/s, 3.05 s to first answer token) are AA's measurements of the single reasoning variant it lists, which carries no effort label even though the API exposes low/medium/high — hence effort null. Pricing is $1.00 input on cache miss / $0.05 on cache hit / $2.70 output including reasoning tokens; max output 64K per StepFun's platform docs. Weights are promised for 2026-10-15 in BF16 (~1.2 TB) with no licence named, so openWeights stays false until they ship. No arena.ai listing, no ARC Prize entry surfaced, knowledge cutoff undisclosed. predecessorId is null: StepFun frames it as a new flagship, not a replacement for Step 3.7 Flash.

For developers

API model strings

  • stepfunstep-5-preview

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News