StepFun exposes StepAudio 2.5 Realtime through model ID step-2.5-realtime. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-29. Next refresh: weekly.
Quick Start
- 1
- 2Use the StepFun SDK or REST API to call
step-2.5-realtime— see the documentation for request format.
Code Examples
Pricing on StepFun
Capabilities
About StepAudio 2.5 Realtime
StepAudio 2.5 Realtime is StepFun's end-to-end real-time conversational voice model. It handles speech input and produces speech output through a single unified architecture with no intermediate ASR/TTS pipeline steps. Key capabilities include persona-consistent roleplay via dedicated RLHF training on million-scale persona data, paralinguistic comprehension (detecting and responding to tone, emotion, and speaking rate), and low-latency dialogue. Supports Chinese and English. Available via WebSocket API (step-2.5-realtime). Analogous in function to OpenAI's GPT Realtime models.