StepAudio 2.5 ASR

Released
2026-05-22
Last refreshed
2026-06-29
Status
Researched 135d ago
ProprietaryCommercial use: conditionalMultimodalVision
StepFun releases · 7 in the last 12 months · this family litChangelog →
Specifications
Released
2026-05-22
Parameters
4B
Specialization
speech-recognition
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

One of China's leading AI 'Six Tigers'.

Shanghai, China
Founded 2023
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 1 route · StepFun

About

StepAudio 2.5 ASR is StepFun's automatic speech recognition model. At 4B parameters, it introduces Multi-Token Prediction (MTP) technology to parallelly predict multiple tokens per decoding step, enabling transcription of 5 minutes of audio in approximately 1 second. Achieves 400% higher throughput and 60% lower latency compared to prior StepFun ASR systems while maintaining state-of-the-art accuracy. Supports Chinese and English; accepts PCM, OGG, MP3, and WAV formats. Available via the StepFun API (model: stepaudio-2.5-asr). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
StepFun--
ServerlessPartial

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.