StepAudio 2.5 Models by StepFun
Last refreshed 2026-06-29. Next refresh: weekly.
Details
Capabilities
Links
WebsiteAbout
StepAudio 2.5 is StepFun's unified audio-language foundation model family, introduced in May 2026 (arXiv:2605.23463). It covers three API-accessible capabilities — text-to-speech (TTS), automatic speech recognition (ASR), and real-time conversational voice (Realtime) — all built on a shared decoder architecture. The family claims top scores across five voice AI benchmarks, surpassing GPT Realtime and Gemini Live on tested dimensions. Supports Chinese and English.
Decision facts
- Best fit
- voicespeech recognitiontext to speech
- Capability starting point
- StepAudio 2.5 Realtime with multimodal inputs
- Lowest tracked input
- Not tracked
- Closest related family
- Step
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs voice, multimodal inputs, and audio.
Use when the workload needs speech recognition, 4B parameters, and multimodal inputs.
Use when the workload needs text to speech, multimodal inputs, and audio.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| StepAudio 2.5 Realtime | Use when the workload needs voice, multimodal inputs, and audio. | 2026-05 | voicemultimodal inputsaudio | Current |
| StepAudio 2.5 ASR | Use when the workload needs speech recognition, 4B parameters, and multimodal inputs. | 2026-05 | speech recognition4B parametersmultimodal inputs | Current |
| StepAudio 2.5 TTS | Use when the workload needs text to speech, multimodal inputs, and audio. | 2026-04 | text to speechmultimodal inputsaudio | Current |
Release Timeline
2 release groupsSpecifications(3 models)
| Model | Released | Parameters | Multimodal |
|---|---|---|---|
| StepAudio 2.5 Realtime | 2026-05 | — | Yes |
| StepAudio 2.5 ASR | 2026-05 | 4B | Yes |
| StepAudio 2.5 TTS | 2026-04 | — | Yes |
