MAI-Voice-2.1

Released
2026-10-01
Last refreshed
2026-10-06
Status
Researched today
ProprietaryCommercial use: conditionalAudio
Microsoft AI releases · 16 in the last 12 monthsChangelog →
Specifications
Family
MAI
Released
2026-10-01
Architecture
Audio / Speech
Specialization
text-to-speech
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Training
Fine-tuned
Created by

Applied AI products and platforms from Microsoft

Redmond, Washington, United States
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 2 routes · Microsoft Foundry

About

MAI-Voice-2.1 is Microsoft AI's highest-fidelity text-to-speech model, launched 1 October 2026. It supports 23 languages and 26 locales with a single voice speaking all of them with a native accent, voice cloning from a few seconds of reference audio, and strong long-form generation. Priced at $22 per 1M characters.

Capabilities

Audio

Benchmark peer barsfor Coding

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

MAI-Voice-2.1