GLM-4-Voice
Last refreshed
2026-07-11
Status
Researched 45d ago
ProprietaryCommercial use: conditionalMultimodalVisionAudio
GLM-4-Voice is worth evaluating for vision when its provider route and context window match the workload.
Use it for
- Teams evaluating vision
- Buyers comparing 1 tracked provider route
Do not use it for
- Strict JSON or tool-calling flows
Specifications
- Family
- GLM-4
- Architecture
- Audio / Speech
- Specialization
- realtime-voice
- Openness
- Proprietary
- License
- ProprietaryCommercial use: conditional
- Weights
- Not released
- Code
- Unknown
Created by
Pricing
Output / 1M
-
Input / 1M
-
Cheapest of 1 route · Zhipu AI GLM API
Links
About
GLM-4-Voice is Zhipu AI's end-to-end speech model for audio or text input and audio output through a distinct voice API surface.
Top use-case fit
Vision
Included by capability and metadata signals in the decision map.
Provider price ladder
Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.
| Provider | Input / 1M | Output / 1M | Route |
|---|---|---|---|
| Zhipu AI GLM API | - | - | ServerlessPartial |
Capabilities
MultimodalAudio
Benchmark peer barsfor Vision
No task-mapped benchmark peers are available for this model yet.
Migration checks
No linked migration route is available for this model yet.
Created by
Pricing
Output / 1M
-
Input / 1M
-
Cheapest of 1 route · Zhipu AI GLM API
Links