25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / StepFun / Step / Step 5 Preview

Step 5 Preview

StepFun's next-generation flagship base model for coding and professional knowledge work, with native text, image, and video input and a 1M-token context window

Line

Step

Weights

API only

Released

2026-09-16

Context

1M

Input

$0.959 / 1M

Output

$2.70 / 1M

Max output

66K

Coverage

3 signals

Overview

Step 5 Preview is StepFun's next-generation flagship base model for coding and professional knowledge work. It accepts native text, image, and video input and returns text. It is the latest in the step line, succeeding Step 3.7 Flash. It has a 1,000,000-token context window and a maximum output of 65,536 tokens. It supports reasoning mode and tool calling. Weights are not open. Pricing starts at $0.959 per million input tokens and $2.70 per million output tokens, with cached input at $0.048 per million tokens.

Specs

Released2026-09-16
LineStepFun · Step
WeightsAPI only
Context1M tokens
Max output66K tokens
Inputimage, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idstep-5-preview

Pricing

Input$0.959 / 1M tokens
Cached input$0.048 / 1M tokens
Output$2.70 / 1M tokens

Against Step 3.7 Flash: window 256K to 1M, input $0.185 to $0.959, output $1.11 to $2.70, cached input $0.03 to $0.048.

Range across the hosts that serve it: up to $1.08 per 1M input tokens.

Strengths

  • 1M-token context window for long documents
  • Native image, text, and video input
  • Reasoning mode and tool calling support
  • Max output of 65,536 tokens
  • Flagship for coding and professional knowledge work

Best for

  • Reach for it for coding tasks that need long context
  • Reach for it for professional knowledge work with multimodal input
  • Reach for it for reasoning-heavy tasks with tool calling
  • Reach for it for processing large documents or videos

How to access

ProviderModel id
AIHubMixstep-5-preview
Eden AIstep-5-preview
EmpirioLabs AIstep-5-preview
NanoGPTstep-5-preview
StepFun (China)step-5-preview
StepFun (Global)step-5-preview
StepFun Step Plan (China)step-5-preview
StepFun Step Plan (Global)step-5-preview
Vercel AI Gatewaystep-5-preview

9 hosts serve this model at the price above · the maker's documentation

What the digest said

Step: every version

VersionReleasedContextInput
Step 5 PreviewCurrent2026-09-161M$0.959
Step 3.7 Flash2026-05-29256K$0.185
Step 3.5 Flash 26032026-04-02256K$0.10
Step 3.5 Flash2026-01-29256K$0.09

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and max output?
Context window is 1,000,000 tokens. Max output is 65,536 tokens.
Does it support image and video input?
Yes, it accepts image, text, and video input, and returns text.
What is the pricing?
Input from $0.959 per million tokens, output $2.70 per million, cached input $0.048 per million.
ProprietaryReasoning1M context