Using GLM-5 on NVIDIA NIM

Implementation guide · GLM-5 · Zhipu AI

ServerlessOpen Source

NVIDIA NIM exposes GLM-5 through model ID z-ai/glm5. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-30. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at NVIDIA NIM and generate an API key.
  2. 2
    Use the NVIDIA NIM SDK or REST API to call z-ai/glm5 — see the documentation for request format.
  3. 3
    You'll be billed . See full pricing.

Code Examples

See NVIDIA NIM documentation for integration details.

Pricing on NVIDIA NIM

Capabilities

ReasoningJSON / Tool useStructured OutputsPrompt Caching

About GLM-5

Flagship open-weight foundation model from Zhipu AI with 744B parameters (40B active per token) in Mixture of Experts architecture. Trained on 28.5T tokens using DeepSeek Sparse Attention on Huawei Ascend hardware. Achieves state-of-the-art performance on coding and agentic benchmarks (SWE-bench Verified: 77.8%). Supports autonomous planning, multi-step tool use, and self-correction.

Model Specs

Released2026-02-11
Parameters744B total, 40B active
Context200k
ArchitectureMixture of Experts
Knowledge cutoff2025-11

Provider

NVIDIA NIM

NVIDIA

Santa Clara, California, United States