GLM-4-Voice
Last refreshed
2026-07-11
Status
Researched 91d ago
ProprietaryCommercial use: conditionalMultimodalVisionAudio
Zhipu AI releases · 14 in the last 12 months · this family litChangelog →
Specifications
- Family
- GLM-4
- Architecture
- Audio / Speech
- Specialization
- realtime-voice
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
Created by
Pricing
Output / 1M
-
Input / 1M
-
Cheapest of 1 route · Zhipu AI GLM API
Links
About
GLM-4-Voice is Zhipu AI's end-to-end speech model for audio or text input and audio output through a distinct voice API surface.
Provider price ladder
Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| Zhipu AI GLM API | - | - | ServerlessPartial |
Capabilities
MultimodalAudio
Benchmark peer barsfor Vision
No task-mapped benchmark peers are available for this model yet.
Migration checks
No linked migration route is available for this model yet.
Created by
Pricing
Output / 1M
-
Input / 1M
-
Cheapest of 1 route · Zhipu AI GLM API
Links