LLM Reference

StepAudio 2.5 ASR

Released
2026-05-22
Last refreshed
2026-06-29
Status
Researched 100d ago
ProprietaryCommercial use: conditionalMultimodalVision

StepAudio 2.5 ASR is worth evaluating for vision when its provider route and context window match the workload.

Use it for

  • Teams evaluating vision
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Strict JSON or tool-calling flows
Specifications
Released
2026-05-22
Parameters
4B
Specialization
speech-recognition
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

One of China's leading AI 'Six Tigers'.

Shanghai, China
Founded 2023
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 1 route · StepFun

About

StepAudio 2.5 ASR is StepFun's automatic speech recognition model. At 4B parameters, it introduces Multi-Token Prediction (MTP) technology to parallelly predict multiple tokens per decoding step, enabling transcription of 5 minutes of audio in approximately 1 second. Achieves 400% higher throughput and 60% lower latency compared to prior StepFun ASR systems while maintaining state-of-the-art accuracy. Supports Chinese and English; accepts PCM, OGG, MP3, and WAV formats. Available via the StepFun API (model: stepaudio-2.5-asr). Part of the unified StepAudio 2.5 architecture described in arXiv:2605.23463.

Top use-case fit

Vision

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
StepFun--
ServerlessPartial

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.