GLM-TTS Models by Zhipu AI
Zhipu AI releases · 14 in the last 12 monthsChangelog →
2 models
Last refreshed 2026-07-11. Next refresh: weekly.
Capabilities
MultimodalAll models
Links
WebsiteAbout
GLM-TTS is Zhipu AI's text-to-speech family for streaming and non-streaming synthesis, including its named reusable voice-cloning model surface.
Decision facts
- Best fit
- audiotext to speechvision and multimodal work
- Capability starting point
- GLM-TTS with multimodal inputs
- Lowest tracked input
- Not tracked
- Closest related family
- AssemblyAI
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
2 in view
GLM-TTSCurrent
Use when the workload needs text to speech, multimodal inputs, and audio.
Unknown releasetext to speechmultimodal inputsaudio
GLM-TTS-CloneCurrent
Use when the workload needs text to speech, multimodal inputs, and audio.
Unknown releasetext to speechmultimodal inputsaudio
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| GLM-TTS | Use when the workload needs text to speech, multimodal inputs, and audio. | Unknown release | text to speechmultimodal inputsaudio | Current |
| GLM-TTS-Clone | Use when the workload needs text to speech, multimodal inputs, and audio. | Unknown release | text to speechmultimodal inputsaudio | Current |
Release Timeline
1 release groupUnknown release
2 current
GLM-TTS
Currenttext to speechmultimodal inputsaudio
GLM-TTS-Clone
Currenttext to speechmultimodal inputsaudio
Specifications(2 models)
| Model | Released | Multimodal |
|---|---|---|
| GLM-TTS | — | Yes |
| GLM-TTS-Clone | — | Yes |

