StepFun exposes StepAudio 2.5 TTS through model ID step-audio-2.5-tts. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-29. Next refresh: weekly.
Quick Start
- 1
- 2Use the StepFun SDK or REST API to call
step-audio-2.5-tts— see the documentation for request format.
Code Examples
Pricing on StepFun
Capabilities
About StepAudio 2.5 TTS
StepAudio 2.5 TTS is StepFun's contextual text-to-speech model with fine-grained expressive control. Unlike tag-based TTS systems, it accepts plain natural language instructions to control emotion, pacing, pauses, and delivery. Supports zero-shot voice cloning with full timbre and emotion control. Priced at $0.85 per 10,000 characters (input text). Supports Chinese and English. Available via StepFun API (model: step-audio-2.5-tts). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.