LLM Reference

GLM-TTS Models by Zhipu AI

Zhipu AIProprietaryAudio
2 models

Last refreshed 2026-07-11. Next refresh: weekly.

Details

ResearcherZhipu AI
LicenseProprietary
Commercial useCommercial use: conditional
Models2

Capabilities

MultimodalAll models

Links

Website

About

GLM-TTS is Zhipu AI's text-to-speech family for streaming and non-streaming synthesis, including its named reusable voice-cloning model surface.

Decision facts

Best fit
audiotext to speechvision and multimodal work
Capability starting point
GLM-TTS with multimodal inputs
Lowest tracked input
Not tracked
Closest related family
GLM-5

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

2 in view
GLM-TTSCurrent

Use when the workload needs text to speech, multimodal inputs, and audio.

Unknown releasetext to speechmultimodal inputsaudio

Use when the workload needs text to speech, multimodal inputs, and audio.

Unknown releasetext to speechmultimodal inputsaudio

Release Timeline

1 release group
Unknown release
2 current
GLM-TTS
text to speechmultimodal inputsaudio
Current
GLM-TTS-Clone
text to speechmultimodal inputsaudio
Current

Specifications(2 models)

GLM-TTS model specifications comparison
ModelReleasedMultimodal
GLM-TTSYes
GLM-TTS-CloneYes

Available From(1 provider)