token.app › Zhipu AI › GLM 5.2
GLM 5.2 pricing & benchmarks
GLM 5.2 is available from Zhipu AI at $0.252 per million input tokens and $0.792 per million output tokens ($0.387 blended at 3:1). It accepts up to 1,048,576 tokens of context. GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Pricing
Per million tokens · OpenRouter
Input$0.252
Output$0.792
Blended (3:1)$0.387
Specification
z-ai/glm-5.2
Context window1.0M
Max output128K
ReleasedJun 2026
LicenceOpen weights
Modalitiesin:text, out:text
Benchmarks
Independently run — not vendor self-reports. Every score links its source.
| Benchmark | Score | Config | Run | Source |
|---|---|---|---|---|
| GPQA Diamond | 91.9% ±1.6 | max effort | 2026-06-24 | Independently run · Epoch AI |
| SWE-bench Verified | 78.7% ±1.9 | max effort | 2026-06-25 | Independently run · Epoch AI |
| FrontierMath T1–3 | 59.2% ±3.0 | max effort | 2026-06-19 | Independently run · Epoch AI |
| FrontierMath T4 | 29.3% ±7.2 | max effort | 2026-06-19 | Independently run · Epoch AI |
| SimpleQA Verified | 38.1% ±1.5 | max effort | 2026-06-19 | Independently run · Epoch AI |
| OTIS Mock AIME | 86.4% ±4.2 | max effort | 2026-06-25 | Independently run · Epoch AI |
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Where to run GLM 5.2
32 options across 27 providers · 9.2× between cheapest and dearest. Blended at 3:1. ⚠ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| StreamLake | $0.387 | $0.252 / $0.792 | 1.0M | fp8 | 99.5% |
| Baidu | $0.409 | $0.266 / $0.836 | 1.0M | fp8 | 99.9% |
| Novita | $0.447 | $0.291 / $0.915 | 1.0M | fp8 | 98.5% |
| Decart | $0.990 | $0.720 / $1.80 | 1.0M | fp4 | 99.9% |
| DeepInfra | $1.16 | $0.750 / $2.40 | 1.0M | fp4 | 97.9% |
| CoreWeave ⚠ 262K context, not 1.0M |
$1.18 | $0.760 / $2.42 | 262K | fp4 | 99.9% |
| AkashML ⚠ 97K context, not 1.0M |
$1.18 | $0.770 / $2.42 | 97K | fp8 | 99.5% |
| Inceptron | $1.29 | $0.750 / $2.90 | 1.0M | fp4 | 98.5% |
| GMICloud | $1.42 | $0.924 / $2.90 | 1.0M | fp8 | 99.5% |
| Sail Research | $1.46 | $0.900 / $3.15 | 1.0M | fp8 | 99.8% |
| Alibaba | $1.48 | $0.966 / $3.04 | 1.0M | fp8 | 99.7% |
| Phala | $1.74 | $1.13 / $3.56 | 1.0M | — | 96.5% |
| SiliconFlow | $1.83 | $1.19 / $3.74 | 1.0M | fp8 | 99.2% |
| Morph | $1.85 | $1.10 / $4.10 | 1.0M | — | 96.8% |
| Ambient ⚠ 203K context, not 1.0M |
$1.89 | $1.05 / $4.40 | 203K | fp8 | 94.4% |
| Wafer | $1.94 | $1.26 / $3.96 | 1.0M | fp4 | 99.5% |
| AtlasCloud | $1.94 | $1.26 / $3.96 | 1.0M | fp8 | 99.7% |
| Z.AI | $2.15 | $1.40 / $4.40 | 1.0M | fp8 | 99.7% |
| Fireworks | $2.15 | $1.40 / $4.40 | 1.0M | — | 98.4% |
| Cloudflare ⚠ 262K context, not 1.0M |
$2.15 | $1.40 / $4.40 | 262K | — | 100.0% |
| Friendli | $2.15 | $1.40 / $4.40 | 1.0M | — | 99.0% |
| Parasail ⚠ 262K context, not 1.0M |
$2.15 | $1.40 / $4.40 | 262K | fp4 | 96.4% |
| Venice | $2.15 | $1.40 / $4.40 | 1M | fp8 | 98.9% |
| Together ⚠ 512K context, not 1.0M |
$2.15 | $1.40 / $4.40 | 512K | — | 91.3% |
| Crusoe | $2.15 | $1.40 / $4.40 | 1.0M | fp8 | 96.2% |
| DigitalOcean ⚠ 262K context, not 1.0M |
$2.15 | $1.40 / $4.40 | 262K | — | 99.7% |
| BaseTen | $2.15 | $1.40 / $4.40 | 1.0M | fp8 | 99.9% |
| Wafer (fast) | $3.22 | $2.10 / $6.60 | 1.0M | fp4 | 99.7% |
| Fireworks (fast) | $3.22 | $2.10 / $6.60 | 1.0M | — | 92.1% |
| Cloudflare (fast) ⚠ 262K context, not 1.0M |
$3.22 | $2.10 / $6.60 | 262K | — | 99.8% |
| BaseTen (fast) ⚠ 524K context, not 1.0M |
$3.22 | $2.10 / $6.60 | 524K | fp8 | 99.9% |
| Alibaba (fast) | $3.55 | $2.31 / $7.26 | 1.0M | fp8 | 99.8% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 — the score above is not host-specific.
Compare GLM 5.2
FAQ
How much does GLM 5.2 cost?
GLM 5.2 costs $0.252 per million input tokens and $0.792 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.387 per million tokens.
What is GLM 5.2's context window?
GLM 5.2 accepts up to 1,048,576 tokens of context, and can generate up to 128,000 output tokens.
How good is GLM 5.2 on benchmarks?
GLM 5.2 scores 91.9% on GPQA Diamond, independently run and published by Epoch AI. 6 benchmark results are listed on this page, each with its run date and source.