Azure Speech Services Models by Microsoft Research
2 models2023
Last refreshed 2026-06-07. Next refresh: weekly.
Details
ResearcherMicrosoft Research
LicenseProprietary
Commercial useCommercial use: conditional
Models2
Released2023
Capabilities
Multimodal1 of 2 models
Links
WebsiteAbout
Microsoft Azure Speech Services model family for speech-to-text and text-to-speech APIs.
Decision facts
- Best fit
- audiotext to speechspeech recognition
- Capability starting point
- Azure Speech Services (STT) with multimodal inputs
- Lowest tracked input
- Not tracked
- Closest related family
- Harrier
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
2 in view
Azure Speech Services (TTS)Current
Use when the workload needs text to speech and audio.
2023-01text to speechaudio
Azure Speech Services (STT)Current
Use when the workload needs speech recognition, multimodal inputs, and audio.
2023-01speech recognitionmultimodal inputsaudio
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Azure Speech Services (TTS) | Use when the workload needs text to speech and audio. | 2023-01 | text to speechaudio | Current |
| Azure Speech Services (STT) | Use when the workload needs speech recognition, multimodal inputs, and audio. | 2023-01 | speech recognitionmultimodal inputsaudio | Current |
Release Timeline
1 release group2023-01
2 current
Azure Speech Services (STT)
Currentspeech recognitionmultimodal inputsaudio
Azure Speech Services (TTS)
Currenttext to speechaudio
Specifications(2 models)
| Model | Released | Multimodal |
|---|---|---|
| Azure Speech Services (TTS) | 2023-01 | No |
| Azure Speech Services (STT) | 2023-01 | Yes |
