MAI Models by Microsoft AI
Last refreshed 2026-09-04. Next refresh: weekly.
Details
Capabilities
Links
WebsiteAbout
Microsoft AI (MAI) is Microsoft's proprietary model family for Copilot and Azure AI Foundry. The lineup now spans reasoning, coding, image generation/editing, speech synthesis, and transcription models, including MAI-Thinking-1, MAI-Code-1, MAI-Code-1-Flash, MAI-Image-2.5, MAI-Image-2.5-Flash, MAI-Voice-2, and MAI-Transcribe-1.5 alongside the earlier MAI image and speech releases.
Decision facts
- Best fit
- image generationspeech recognitionaudio
- Capability starting point
- MAI-Thinking-1 with 256k context and reasoning and JSON / Tool use
- Lowest tracked input
- MAI-Transcribe-1 · $0.36/1M · Microsoft Foundry
- Closest related family
- Claude 3
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs image generation, 32k context, and multimodal inputs.
Use when the workload needs speech recognition, multimodal inputs, and audio.
Use when the workload needs reasoning, 256k context, and JSON / Tool use.
Use when the workload needs code, 256k context, and reasoning.
Use when the workload needs image generation, 32k context, and multimodal inputs.
Use when the workload needs image generation, 32k context, and multimodal inputs.
Use when the workload needs speech recognition, multimodal inputs, and audio.
Use when the workload needs image generation, 33k context, and multimodal inputs.
Use when the workload needs speech recognition, multimodal inputs, and audio.
Use when the workload needs text to speech, multimodal inputs, and audio.
Use when the workload needs image generation and multimodal inputs.
Use when the workload needs reasoning, 164k context, and 671B parameters.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| MAI-Image-2.6-Flash | Use when the workload needs image generation, 32k context, and multimodal inputs. | 2026-09 | image generation32k contextmultimodal inputs | Current |
| MAI-Transcribe-2 | Use when the workload needs speech recognition, multimodal inputs, and audio. | 2026-09 | speech recognitionmultimodal inputsaudio | Current |
| MAI-Thinking-1 | Use when the workload needs reasoning, 256k context, and JSON / Tool use. | 2026-06 | reasoning256k contextJSON / Tool use | Current |
| MAI-Code-1 | Use when the workload needs code. | 2026-06 | code | Current |
| MAI-Code-1-Flash | Use when the workload needs code, 256k context, and reasoning. | 2026-06 | code256k contextreasoning | Current |
| MAI-Image-2.5 | Use when the workload needs image generation, 32k context, and multimodal inputs. | 2026-06 | image generation32k contextmultimodal inputs | Current |
| MAI-Image-2.5-Flash | Use when the workload needs image generation, 32k context, and multimodal inputs. | 2026-06 | image generation32k contextmultimodal inputs | Current |
| MAI-Voice-2 | Use when the workload needs text to speech and audio. | 2026-06 | text to speechaudio | Current |
| MAI-Transcribe-1.5 | Use when the workload needs speech recognition, multimodal inputs, and audio. | 2026-06 | speech recognitionmultimodal inputsaudio | Current |
| MAI-Image-2e | Use when the workload needs image generation, 33k context, and multimodal inputs. | 2026-04 | image generation33k contextmultimodal inputs | Current |
| MAI-Transcribe-1 | Use when the workload needs speech recognition, multimodal inputs, and audio. | 2026-04 | speech recognitionmultimodal inputsaudio | Current |
| MAI-Voice-1 | Use when the workload needs text to speech, multimodal inputs, and audio. | 2026-04 | text to speechmultimodal inputsaudio | Current |
| MAI-Image-2 | Use when the workload needs image generation and multimodal inputs. | 2026-03 | image generationmultimodal inputs | Current |
| MAI-DS-R1 | Use when the workload needs reasoning, 164k context, and 671B parameters. | 2025-04 | reasoning164k context671B parameters | Current |
Release Timeline
5 release groupsSpecifications(14 models)
| Model | Released | Context | Parameters | Vision | Multimodal | Reasoning | JSON / Tool use |
|---|---|---|---|---|---|---|---|
| MAI-Image-2.6-Flash | 2026-09 | 32k | — | Yes | Yes | No | No |
| MAI-Transcribe-2 | 2026-09 | — | — | No | Yes | No | No |
| MAI-Thinking-1 | 2026-06 | 256k | 1T total / 35B active | No | No | Yes | Yes |
| MAI-Code-1 | 2026-06 | — | — | No | No | No | No |
| MAI-Code-1-Flash | 2026-06 | 256k | — | No | No | Yes | Yes |
| MAI-Image-2.5 | 2026-06 | 32k | — | Yes | Yes | No | No |
| MAI-Image-2.5-Flash | 2026-06 | 32k | — | Yes | Yes | No | No |
| MAI-Voice-2 | 2026-06 | — | — | No | No | No | No |
| MAI-Transcribe-1.5 | 2026-06 | — | — | No | Yes | No | No |
| MAI-Image-2e | 2026-04 | 33k | — | Yes | Yes | No | No |
| MAI-Transcribe-1 | 2026-04 | — | — | No | Yes | No | No |
| MAI-Voice-1 | 2026-04 | — | — | No | Yes | No | No |
| MAI-Image-2 | 2026-03 | — | — | Yes | Yes | No | No |
| MAI-DS-R1 | 2025-04 | 164k | 671B | No | No | Yes | No |
Available From(1 provider)
Pricing
| Model | Provider | Input / 1M | Output / 1M | Type |
|---|---|---|---|---|
| MAI-Transcribe-1 | Microsoft Foundry | $0.36 | — | Serverless |
| MAI-Code-1-Flash | Microsoft Foundry | $0.75 | $4.5 | Serverless |
| MAI-Image-2 | Microsoft Foundry | $5 | $33 | Serverless |
| MAI-Voice-1 | Microsoft Foundry | $22 | — | Serverless |
Comparisons
- MAI-Thinking-1 vs Claude Opus 4.6
- MAI-Thinking-1 vs Claude Sonnet 4.6
- Claude Fable 5 vs MAI-Thinking-1
- MAI-Code-1-Flash vs Claude Haiku 4.5
Models(14)
MAI-Image-2.6-Flash
MAI-Transcribe-2
MAI-Thinking-1
MAI-Code-1
MAI-Code-1-Flash
MAI-Image-2.5
MAI-Image-2.5-Flash
MAI-Voice-2
MAI-Transcribe-1.5
MAI-Image-2e
MAI-Transcribe-1
MAI-Voice-1
MAI-Image-2
MAI-DS-R1
