StepAudio 2.5 TTS
- Family
- StepAudio 2.5
- Released
- 2026-04-16
- Specialization
- text-to-speech
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
Cheapest of 1 route · StepFun
About
StepAudio 2.5 TTS is StepFun's contextual text-to-speech model with fine-grained expressive control. Unlike tag-based TTS systems, it accepts plain natural language instructions to control emotion, pacing, pauses, and delivery. Supports zero-shot voice cloning with full timbre and emotion control. Priced at $0.85 per 10,000 characters (input text). Supports Chinese and English. Available via StepFun API (model: step-audio-2.5-tts). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.
Provider price ladder
Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| StepFun | - | - | ServerlessPartial |
Capabilities
Benchmark peer barsfor Vision
No task-mapped benchmark peers are available for this model yet.
Benchmark scores(1)
| Benchmark | Score | Version | Evaluation | Source |
|---|---|---|---|---|
| Artificial Analysis TTS Arena ELO | 1187.0 | arena-eloObserved 2026-05-24 | — | Source |
Migration checks
No linked migration route is available for this model yet.
Cheapest of 1 route · StepFun