Using StepAudio 2.5 Realtime on StepFun

Implementation guide · StepAudio 2.5 · StepFun

Serverless

StepFun exposes StepAudio 2.5 Realtime through model ID step-2.5-realtime. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-29. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at StepFun and generate an API key.
  2. 2
    Use the StepFun SDK or REST API to call step-2.5-realtime — see the documentation for request format.

Code Examples

See StepFun documentation for integration details.

Pricing on StepFun

Capabilities

MultimodalAudio

About StepAudio 2.5 Realtime

StepAudio 2.5 Realtime is StepFun's end-to-end real-time conversational voice model. It handles speech input and produces speech output through a single unified architecture with no intermediate ASR/TTS pipeline steps. Key capabilities include persona-consistent roleplay via dedicated RLHF training on million-scale persona data, paralinguistic comprehension (detecting and responding to tone, emotion, and speaking rate), and low-latency dialogue. Supports Chinese and English. Available via WebSocket API (step-2.5-realtime). Analogous in function to OpenAI's GPT Realtime models.

Model Specs

Released2026-05-24

Provider

StepFun