MAI-Voice-2.1-Flash

Released
2026-10-01
Last refreshed
2026-10-06
Status
Researched today
ProprietaryCommercial use: conditionalAudio
Microsoft AI releases · 16 in the last 12 monthsChangelog →
Specifications
Family
MAI
Released
2026-10-01
Architecture
Audio / Speech
Specialization
text-to-speech
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Training
Fine-tuned
Created by

Applied AI products and platforms from Microsoft

Redmond, Washington, United States
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 2 routes · Microsoft Foundry

About

MAI-Voice-2.1-Flash is Microsoft AI's low-latency text-to-speech model for high-volume voice agents, launched 1 October 2026 alongside MAI-Voice-2.1. It covers the same 23 languages and cross-language voices with about 150 ms end-to-end latency and supports voice cloning. Priced at $15 per 1M characters.

Capabilities

Audio

Benchmark peer barsfor Coding

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

MAI-Voice-2.1-Flash