GLM-TTS Models by Zhipu AI
2 models
Last refreshed 2026-07-11. Next refresh: weekly.
Capabilities
MultimodalAll models
Links
WebsiteAbout
GLM-TTS is Zhipu AI's text-to-speech family for streaming and non-streaming synthesis, including its named reusable voice-cloning model surface.
Decision facts
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
2 in view
GLM-TTSCurrent
Use when the workload needs text to speech, multimodal inputs, and audio.
Unknown releasetext to speechmultimodal inputsaudio
GLM-TTS-CloneCurrent
Use when the workload needs text to speech, multimodal inputs, and audio.
Unknown releasetext to speechmultimodal inputsaudio
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| GLM-TTS | Use when the workload needs text to speech, multimodal inputs, and audio. | Unknown release | text to speechmultimodal inputsaudio | Current |
| GLM-TTS-Clone | Use when the workload needs text to speech, multimodal inputs, and audio. | Unknown release | text to speechmultimodal inputsaudio | Current |
Release Timeline
1 release groupUnknown release
2 current
GLM-TTS
Currenttext to speechmultimodal inputsaudio
GLM-TTS-Clone
Currenttext to speechmultimodal inputsaudio
Specifications(2 models)
| Model | Released | Multimodal |
|---|---|---|
| GLM-TTS | — | Yes |
| GLM-TTS-Clone | — | Yes |





