Microsoft Foundry exposes MAI-Voice-2.1 through model ID MAI-Voice-2.1. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-10-06. Next refresh: weekly.
Quick Start
- 1
- 2Use the Microsoft Foundry SDK or REST API to call
MAI-Voice-2.1— see the documentation for request format.
Code Examples
Pricing on Microsoft Foundry
Capabilities
Audio
About MAI-Voice-2.1
MAI-Voice-2.1 is Microsoft AI's highest-fidelity text-to-speech model, launched 1 October 2026. It supports 23 languages and 26 locales with a single voice speaking all of them with a native accent, voice cloning from a few seconds of reference audio, and strong long-form generation. Priced at $22 per 1M characters.
Model Specs
Released2026-10-01
ArchitectureAudio / Speech