LLM Reference

Azure Speech Services Models by Microsoft Research

Microsoft ResearchProprietaryAudio
2 models2023

Last refreshed 2026-06-07. Next refresh: weekly.

Details

LicenseProprietary
Commercial useCommercial use: conditional
Models2
Released2023

Capabilities

Multimodal1 of 2 models

Links

Website

About

Microsoft Azure Speech Services model family for speech-to-text and text-to-speech APIs.

Decision facts

Best fit
audiotext to speechspeech recognition
Capability starting point
Azure Speech Services (STT) with multimodal inputs
Lowest tracked input
Not tracked
Closest related family
Harrier

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

2 in view

Use when the workload needs text to speech and audio.

2023-01text to speechaudio

Use when the workload needs speech recognition, multimodal inputs, and audio.

2023-01speech recognitionmultimodal inputsaudio

Release Timeline

1 release group
2023-01
2 current
Azure Speech Services (STT)
speech recognitionmultimodal inputsaudio
Current
Azure Speech Services (TTS)
text to speechaudio
Current

Specifications(2 models)

Azure Speech Services model specifications comparison
ModelReleasedMultimodal
Azure Speech Services (TTS)2023-01No
Azure Speech Services (STT)2023-01Yes