Last refreshed 2026-06-29. Next refresh: weekly.
Why use StepAudio 2.5 ASR on StepFun?
StepFun offers StepAudio 2.5 ASR with competitive pricing. StepFun is a Chinese AI company providing API access to its Step series of large language and multimodal models.
Setup recipe
Docs fallbackUse the provider REST API or SDKCreate a provider API keymodel: stepaudio-2.5-asrstepaudio-2.5-asrRequest example
stepaudio-2.5-asr.Gotchas
- Use provider model ID "stepaudio-2.5-asr", not the LLMReference slug "step-audio-2-5-asr".
Capabilities
About StepAudio 2.5 ASR
StepAudio 2.5 ASR is StepFun's automatic speech recognition model. At 4B parameters, it introduces Multi-Token Prediction (MTP) technology to parallelly predict multiple tokens per decoding step, enabling transcription of 5 minutes of audio in approximately 1 second. Achieves 400% higher throughput and 60% lower latency compared to prior StepFun ASR systems while maintaining state-of-the-art accuracy. Supports Chinese and English; accepts PCM, OGG, MP3, and WAV formats. Available via the StepFun API (model: stepaudio-2.5-asr). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.