LLM Reference

Step Models by StepFun

StepFunProprietary
11 models2024–2026Up to 256k ctxFrom $0.1/1M input

Last refreshed 2026-06-29. Next refresh: weekly.

Details

ResearcherStepFun
LicenseProprietary
Commercial useCommercial use: conditional
Models11
Released2024–2026
Max context256k

Capabilities

Vision3 of 11 models
Multimodal5 of 11 models
Reasoning2 of 11 models
JSON / Tool use2 of 11 models
Structured Outputs1 of 11 models

About

The Step family of large language and multimodal models from StepFun (阶跃星辰). The series spans proprietary API models and open-weight Flash releases, including Step 3.7 Flash, a 198B-parameter sparse MoE vision-language model with 256K context, Apache 2.0 weights, and selectable reasoning levels for agentic coding, tool use, image, and video workflows.

Decision facts

Best fit
vision and multimodal workreasoningJSON / Tool use
Capability starting point
Step 3.7 Flash with 256k context and reasoning, JSON / Tool use, structured outputs, and multimodal inputs
Lowest tracked input
Step 3.5 Flash · $0.1/1M · OpenRouter
Closest related family
StepAudio 2.5

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

11 in view

Use when the workload needs 256k context, reasoning, and JSON / Tool use.

2026-05256k contextreasoningJSON / Tool use

Use when the workload needs 256k context and reasoning.

2026-01256k contextreasoning

Use when the workload needs 128k context.

2025-10128k context

Use when the workload needs 128k context.

2025-08128k context
Step-2Current

Use when the workload needs 256k context, JSON / Tool use, and multimodal inputs.

2024-09256k contextJSON / Tool usemultimodal inputs

Use when the workload needs multimodal inputs.

2024-07multimodal inputs
Step-1.5VCurrent

Use when the workload needs 128k context and multimodal inputs.

2024-06128k contextmultimodal inputs
Step-1Current

Use when the workload needs 128k context.

2024-04128k context
Step-1VCurrent

Use when the workload needs multimodal inputs.

2024-03multimodal inputs

Use when provider availability and model metadata match the workload.

2024-03
Step-MathCurrent

Use when provider availability and model metadata match the workload.

2024-03

Release Timeline

9 release groups
2026-05
1 current
Step 3.7 Flash
256k contextreasoningJSON / Tool use
Current
2026-01
1 current
Step 3.5 Flash
256k contextreasoning
Current
2025-10
1 current
StepFun Step-2
128k context
Current
2025-08
1 current
StepFun Step-1
128k context
Current
2024-09
1 current
Step-2
256k contextJSON / Tool usemultimodal inputs
Current
2024-07
1 current
Step-1V Turbo
multimodal inputs
Current
2024-06
1 current
Step-1.5V
128k contextmultimodal inputs
Current
2024-04
1 current
Step-1
128k context
Current

Specifications(11 models)

Step model specifications comparison
ModelReleasedContextParametersVisionMultimodalReasoningJSON / Tool useStructured Outputs
Step 3.7 Flash2026-05256k198B (11B active)YesYesYesYesYes
Step 3.5 Flash2026-01256k196B (11B active)NoNoYesNoNo
StepFun Step-22025-10128k1T (MoE)*NoNoNoNoNo
StepFun Step-12025-08128kNoNoNoNoNo
Step-22024-09256k1T (MoE)*YesYesNoYesNo
Step-1V Turbo2024-07NoYesNoNoNo
Step-1.5V2024-06128kYesYesNoNoNo
Step-12024-04128kNoNoNoNoNo
Step-1V2024-03NoYesNoNoNo
Step-Instruct2024-03NoNoNoNoNo
Step-Math2024-03NoNoNoNoNo

Available From(3 providers)

Pricing

Step model pricing by provider
ModelProviderInput / 1MOutput / 1MType
Step 3.5 FlashOpenRouter$0.1$0.3Serverless
Step 3.7 FlashStepFun$0.2$1.15Serverless
Step 3.7 FlashOpenRouter$0.2$1.15Serverless

Popular comparisons in this family