Last refreshed 2026-06-29. Next refresh: weekly.
Why use StepAudio 2.5 TTS on StepFun?
StepFun offers StepAudio 2.5 TTS with competitive pricing. StepFun is a Chinese AI company providing API access to its Step series of large language and multimodal models.
Input / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced
Setup recipe
Docs fallbackInstall
Use the provider REST API or SDKAuth
Create a provider API keyCall
model: step-audio-2.5-ttsModel ID
step-audio-2.5-ttsRequest example
Curated snippets for this provider are not sourced yet. Use StepFun documentation with model ID
step-audio-2.5-tts.Gotchas
- Use provider model ID "step-audio-2.5-tts", not the LLMReference slug "step-audio-2-5-tts".
Capabilities
MultimodalAudio
About StepAudio 2.5 TTS
StepAudio 2.5 TTS is StepFun's contextual text-to-speech model with fine-grained expressive control. Unlike tag-based TTS systems, it accepts plain natural language instructions to control emotion, pacing, pauses, and delivery. Supports zero-shot voice cloning with full timbre and emotion control. Priced at $0.85 per 10,000 characters (input text). Supports Chinese and English. Available via StepFun API (model: step-audio-2.5-tts). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.
Get Started
Model Specs
Released2026-04-16