LLM Reference

GLM-4-Voice

Last refreshed
2026-07-11
Status
Researched 45d ago
ProprietaryCommercial use: conditionalMultimodalVisionAudio

GLM-4-Voice is worth evaluating for vision when its provider route and context window match the workload.

Use it for

  • Teams evaluating vision
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Strict JSON or tool-calling flows
Specifications
Family
GLM-4
Architecture
Audio / Speech
Specialization
realtime-voice
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

Chinese AI research lab developing GLM language models.

Beijing, China
Founded 2019
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 1 route · Zhipu AI GLM API

About

GLM-4-Voice is Zhipu AI's end-to-end speech model for audio or text input and audio output through a distinct voice API surface.

Top use-case fit

Vision

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
Zhipu AI GLM API--
ServerlessPartial

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.