LLM Reference

MAI-Transcribe-2

Released
2026-09-03
Last refreshed
2026-09-03
Status
Researched 1d ago
ProprietaryCommercial use: conditionalMultimodalVisionAudio

MAI-Transcribe-2 is worth evaluating for vision when its provider route and context window match the workload.

Use it for

  • Teams evaluating vision
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Strict JSON or tool-calling flows
Specifications
Family
MAI
Released
2026-09-03
Architecture
Audio / Speech
Specialization
speech-recognition
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Training
Fine-tuned
Created by

Applied AI products and platforms from Microsoft

Redmond, Washington, United States
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 1 route · Microsoft Foundry

About

MAI-Transcribe-2 (Azure Speech enhancedMode.model id MAI-Transcribe-2) is Microsoft AI's third-generation speech-to-text / ASR model, announced September 3, 2026. Specialized transcription API, not a chat LLM. First-party: 60 languages; speaker diarization; word-level timestamps; keyword biasing; verbatim/clean transcription styles; code switching; automatic language identification; noise-robust real-world audio. Microsoft reports first place on FLEURS across 60 languages at 5.2% average WER, Artificial Analysis accuracy-latency Pareto leadership, and second place on the Artificial Analysis WER leaderboard. Available to demo via Microsoft Foundry, MAI Playground, and Open Router; Azure Speech docs mark MAI-Transcribe as public preview. Compare it for Speech recognition.

Top use-case fit

Vision

Included by capability and metadata signals in the decision map.

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

MAI-Transcribe-2