MAI-Transcribe-2-Streaming

Released
2026-10-01
Last refreshed
2026-10-06
Status
Researched today
ProprietaryCommercial use: conditionalMultimodalVisionAudio
Microsoft AI releases · 16 in the last 12 monthsChangelog →
Specifications
Family
MAI
Released
2026-10-01
Architecture
Audio / Speech
Specialization
speech-recognition
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Training
Fine-tuned
Created by

Applied AI products and platforms from Microsoft

Redmond, Washington, United States
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 2 routes · Microsoft Foundry

About

MAI-Transcribe-2-Streaming is Microsoft AI's first streaming speech-to-text model, launched 1 October 2026. Audio streams over a WebSocket and transcripts return incrementally while the speaker talks, with partial results that are then finalised. It is offered at an introductory $0.54 per hour of audio through the end of 2026.

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

MAI-Transcribe-2-Streaming