Model registry / StepFun / Step / Step 5 Preview
Step 5 Preview
StepFun's next-generation flagship base model for coding and professional knowledge work, with native text, image, and video input and a 1M-token context window
Line
Step
Weights
API only
Released
2026-09-16
Context
1M
Input
$0.959 / 1M
Output
$2.70 / 1M
Max output
66K
Coverage
3 signals
Overview
Step 5 Preview is StepFun's next-generation flagship base model for coding and professional knowledge work. It accepts native text, image, and video input and returns text. It is the latest in the step line, succeeding Step 3.7 Flash. It has a 1,000,000-token context window and a maximum output of 65,536 tokens. It supports reasoning mode and tool calling. Weights are not open. Pricing starts at $0.959 per million input tokens and $2.70 per million output tokens, with cached input at $0.048 per million tokens.
Specs
Pricing
Against Step 3.7 Flash: window 256K to 1M, input $0.185 to $0.959, output $1.11 to $2.70, cached input $0.03 to $0.048.
Range across the hosts that serve it: up to $1.08 per 1M input tokens.
Strengths
- 1M-token context window for long documents
- Native image, text, and video input
- Reasoning mode and tool calling support
- Max output of 65,536 tokens
- Flagship for coding and professional knowledge work
Best for
- Reach for it for coding tasks that need long context
- Reach for it for professional knowledge work with multimodal input
- Reach for it for reasoning-heavy tasks with tool calling
- Reach for it for processing large documents or videos
How to access
9 hosts serve this model at the price above · the maker's documentation
What the digest said
- 8Artificial Analysis 评测 Step 5 Preview: Intelligence Index 得 44 分,成本约为同级模型 1/2.82026-09-22
- 9刚刚,阶跃Step 5 Preview发布!一举杀进全球开源前三2026-09-21
- 8开源Top2!实测阶跃Step 5 Preview,真有点猛啊…2026-09-21
3 signals of the archive mention this model, each scored against the published bar. This list is rebuilt from the archive at every build, so it cannot go stale.
Step: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and max output?
- Context window is 1,000,000 tokens. Max output is 65,536 tokens.
- Does it support image and video input?
- Yes, it accepts image, text, and video input, and returns text.
- What is the pricing?
- Input from $0.959 per million tokens, output $2.70 per million, cached input $0.048 per million.