StepAudio 2.5 ASR on StepFun

StepAudio 2.5 · StepFun

Serverless

Last refreshed 2026-06-29. Next refresh: weekly.

Why use StepAudio 2.5 ASR on StepFun?

StepFun offers StepAudio 2.5 ASR with competitive pricing. StepFun is a Chinese AI company providing API access to its Step series of large language and multimodal models.

Input / 1M
-
Output / 1M
-
Cache
Not sourced
Batch
Not sourced

Setup recipe

Docs fallback
Install
Use the provider REST API or SDK
Auth
Create a provider API key
Call
model: stepaudio-2.5-asr
Model ID
stepaudio-2.5-asr

Request example

Curated snippets for this provider are not sourced yet. Use StepFun documentation with model ID stepaudio-2.5-asr.

Gotchas

  • Use provider model ID "stepaudio-2.5-asr", not the LLMReference slug "step-audio-2-5-asr".

Capabilities

MultimodalAudio

About StepAudio 2.5 ASR

StepAudio 2.5 ASR is StepFun's automatic speech recognition model. At 4B parameters, it introduces Multi-Token Prediction (MTP) technology to parallelly predict multiple tokens per decoding step, enabling transcription of 5 minutes of audio in approximately 1 second. Achieves 400% higher throughput and 60% lower latency compared to prior StepFun ASR systems while maintaining state-of-the-art accuracy. Supports Chinese and English; accepts PCM, OGG, MP3, and WAV formats. Available via the StepFun API (model: stepaudio-2.5-asr). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.

Get Started

Model Specs

Released2026-05-22
Parameters4B