StepAudio 2.5 TTS on StepFun

StepAudio 2.5 · StepFun

Serverless

Last refreshed 2026-06-29. Next refresh: weekly.

Why use StepAudio 2.5 TTS on StepFun?

StepFun offers StepAudio 2.5 TTS with competitive pricing. StepFun is a Chinese AI company providing API access to its Step series of large language and multimodal models.

Input / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced

Setup recipe

Docs fallback
Install
Use the provider REST API or SDK
Auth
Create a provider API key
Call
model: step-audio-2.5-tts
Model ID
step-audio-2.5-tts

Request example

Curated snippets for this provider are not sourced yet. Use StepFun documentation with model ID step-audio-2.5-tts.

Gotchas

  • Use provider model ID "step-audio-2.5-tts", not the LLMReference slug "step-audio-2-5-tts".

Capabilities

MultimodalAudio

About StepAudio 2.5 TTS

StepAudio 2.5 TTS is StepFun's contextual text-to-speech model with fine-grained expressive control. Unlike tag-based TTS systems, it accepts plain natural language instructions to control emotion, pacing, pauses, and delivery. Supports zero-shot voice cloning with full timbre and emotion control. Priced at $0.85 per 10,000 characters (input text). Supports Chinese and English. Available via StepFun API (model: step-audio-2.5-tts). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.

Get Started