LLM Reference

GLM-5.2

Released
2026-06-13
Last refreshed
2026-06-24
Status
Researched 54d ago
Open sourceCommercial use: permittedCodingRAGAgentsLong contextClassificationJSON / Tool use

GLM-5.2 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 1m context window
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Vision or document-understanding workloads
Specifications
Family
GLM-5
Released
2026-06-13
Context
1m
Max output
131,072
Parameters
753B total, 40B active
Architecture
Mixture of Experts
Specialization
coding
Openness
Open source
License
MITOSI-approvedCommercial use: permitted
Weights
Available
Code
Unknown
Training
Fine-tuned
Created by

Chinese AI research lab developing GLM language models.

Beijing, China
Founded 2019
Website
Pricing
Output / 1M
$4.40
Input / 1M
$1.40

Cheapest of 1 route · OpenRouter

About

Choose GLM-5.2 when you want an MIT-licensed, self-hostable coding model you can also route cheaply via OpenRouter—not another closed flagship. This page pairs the Coding label with self-reported 62.1% SWE-bench Pro and 82.7% Terminal-Bench 2.1 figures, the open-weights badge, and the cheapest tracked OpenRouter route (live per-token price from seed). Agentic scores—MCP-Atlas 76.8, Tool-Decathlon 48.2—sit under the Agents label. SWE-bench Pro is self-reported on the z.ai Hugging Face card, not an independent evaluation.

GLM-5.2 is an open-source model in the GLM-5 family. The structured metadata tracks a 1m-token context window, reasoning, function calling, tool use, structured outputs, and code execution. This page tracks provider routes through OpenRouter, with the cheapest tracked route listed at $1.4 input and $4.4 output per 1M tokens. Headline tracked benchmarks include SWE-bench Pro 62.1, Google-Proof Q&A 91.2, and AIME 2026 99.2.

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ D

1 relevant benchmark in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
OpenRouter$1.40$4.40
Serverless

Capabilities

ReasoningFunction CallingTool UseStructured OutputsCode Execution

Benchmark peer barsfor Coding

Benchmark scores(11)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
SWE-bench Pro62.1SWE-bench Pro (% resolved, self-reported)Observed 2026-06-13Source
Google-Proof Q&A91.2diamondObserved 2026-06-13Source
AIME 202699.2AIME 2026 (accuracy)Observed 2026-06-13Source
Humanity's Last Exam40.5HLE (accuracy, text-only assumed)Observed 2026-06-13Source
Terminal-Bench 2.182.7Terminal-Bench 2.1 (% tasks completed)Observed 2026-06-13Source
MCP-Atlas76.8MCP-AtlasObserved 2026-06-13Source
Toolathlon48.2Tool-Decathlon (% tasks completed)Observed 2026-06-13Source
CursorBench54.6CursorBench 3.1Observed 2026-06-30
Configuration: GLM 5.2 Max
Harness: CursorBench 3.1
Evaluator: Cursor
Confidence: confirmed
Notes: Highest CursorBench 3.1 score across Cursor's published effort configurations for this base model.
Source
GeneBench-Pro4.6highObserved 2026-06-30Source
CursorBench55.0CursorBench 3.2Observed 2026-07-18
Configuration: GLM 5.2 Max
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $1.76
Tokens/task: 35,946
Steps/task: 58
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source
CursorBench51.5CursorBench 3.2Observed 2026-07-18
Configuration: GLM 5.2 High
Harness: CursorBench 3.2 productized Cursor-agent workflow
Evaluator: Cursor
Confidence: confirmed
Cost/task: $1.19
Tokens/task: 21,829
Steps/task: 49
Notes: Cursor vendor-reported result; not independently reproducible. Results are subject to variance, and small score differences may not be statistically meaningful.
Source

Migration checks

No linked migration route is available for this model yet.

Compare GLM-5.2 with other models

Frequently asked questions

What is the context window of GLM-5.2?

GLM-5.2 has a context window of 1m tokens.

What is the max output of GLM-5.2?

GLM-5.2 can generate up to 131,072 output tokens.

How much does GLM-5.2 cost?

GLM-5.2 is available at $1.40/1M input tokens through OpenRouter.

When was GLM-5.2 released?

GLM-5.2 was released on 2026-06-13.

Which providers offer GLM-5.2?

GLM-5.2 is available from 1 provider: OpenRouter.

What benchmarks has GLM-5.2 been tested on?

GLM-5.2 has been evaluated on 11 benchmarks, including SWE-bench Pro, Google-Proof Q&A, AIME 2026, Humanity's Last Exam, Terminal-Bench 2.1.