Pricing (50% off promo through Sep 9, 2026)
Verified against Z.ai's official pricing page (docs.z.ai). All prices in USD per 1M tokens. Strikethrough = list price.
| Model | Input | Cached Input | Cache Storage | Output | Notes |
|---|---|---|---|---|---|
| GLM-5.3-Flash | Limited-time Free | 50% off. Promo ends 24:00 Sep 9, 2026 (UTC+8, Singapore time) | |||
| GLM-5.3 | $1.40 | $0.26 | Limited-time Free | $4.40 | Reasoning model. ~10× Flash price |
| GLM-5.2 | $1.40 | $0.26 | Limited-time Free | $4.40 | MIT license. Predecessor of Flash |
thinking.type: enabled (cannot be disabled). Thinking tokens are billed as output, so watch costs even on simple tasks. Recommended settings: reasoning_effort: max, temperature: 1, top_p: 0.95.
Primary source: docs.z.ai/guides/overview/pricing (verified 2026-08-31)
Model Specifications
Sources: Z.ai: GLM-5.3-Flash · Official blog · HuggingFace: zai-org/GLM-5.3-Flash
Performance — Beats GLM-5.2, approaches Claude Opus 4.8
Figures below are self-reported by Z.ai's official blog (2026-08-26) and may diverge from independent evaluations.
- Artificial Analysis Intelligence Index v4.1.1: 57 at $0.045/task (discounted) — a level of intelligence Z.ai says previously cost ~10× more.
- DeepSWE v1.1: 63.4 (GLM-5.2 46.2, Claude Opus 4.8 58.0, GPT-5.6 Terra 69.6, Gemini 3.7 Flash 65.3).
- Terminal-Bench 2.1: 84.3 (close to Claude Opus 4.8 85.0, GPT-5.6 Terra 87.4, Gemini 3.7 Flash 85.8).
- AutomationBench v1.0.6: 48.8 (GLM-5.2 26.2) — large agentic gains over GLM-5.2.
- Z.ai Code Bench v1.0: 29.0 at max effort (nearly matching Claude Opus 4.8's 29.5).
- vs GLM-5.2: Outperforms GLM-5.2 across all six coding/agentic benchmarks (often by a wide margin).
- Cost-efficiency: ~1/10 the price of GLM-5.3/5.2 ($1.40/$4.40) with better performance.
Z.ai positions the model as proof that "frontier intelligence does not have to come at frontier cost." Its hybrid sparse + linear attention cuts attention compute by ~3.0× and KV cache by ~4.4× versus GLM-5.3, sharply reducing long-context serving cost. Before release it was tested anonymously as ox-alpha on OpenCode and OpenRouter, quickly becoming that week's most popular model.
Source: Z.ai blog "GLM-5.3-Flash: Frontier Intelligence, Flash Cost" (verified 2026-08-31)
Competitor Price Comparison
Comparison against low-cost Chinese models. USD per 1M tokens (standard text, no cache). GLM-5.3-Flash at promo price.
| Model | Provider | Input | Output | vs Flash | Notes |
|---|---|---|---|---|---|
| GLM-5.3-Flash | Zhipu (Z.ai) | $0.075 | $0.25 | — | Baseline. Promo (thru 9/9). Multimodal |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Flash cheaper | Flash 62.5% cheaper input, 79% cheaper output. Luna text+image only |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Flash cheaper | Flash 70% cheaper input, 83% cheaper output | |
| DeepSeek V4 Flash | DeepSeek | $0.22–$0.44 | $0.66–$1.32 | Flash cheaper for both input and output | Peak/off-peak. Off-peak = half of peak |
| MiniMax M3 | MiniMax | $0.30 | $1.20 | Flash cheaper | After "permanent 50% off" (3rd party) |
| GLM-5.3 / 5.2 | Zhipu (Z.ai) | $1.40 | $4.40 | Flash cheaper | Flash ~1/10 the price. Higher-tier reasoning models |
* GLM-5.3-Flash's $0.075/$0.25 is a promo price (list $0.15/$0.50, thru 9/9). DeepSeek V4 Flash moved to peak/off-peak pricing on 2026-08-16 (cache-miss input $0.22 off-peak / $0.44 peak, output $0.66 off-peak / $1.32 peak). GPT-5.6 Luna comparison is text-input only.
Sources: official pricing pages from each provider (verified 2026-08-31). See full pricing comparison. GPT-5.6 Luna details →
Ideal Use Cases
Low cost × long context × multimodal workloads.
Coding & Agents
Terminal-Bench 84.3, DeepSWE 63.4. Autonomous agents over large codebases.
Visual Coding
Recreate UI/frontends from screenshots, recordings, and URLs with a self-verification loop.
Office & Document Workflows
Generate PPTX/PDF/DOCX/XLSX. Financial research and document processing.
Tool Use & Agents
Toolathlon 78.4. Function calling and structured (JSON) output.
3D & Game Development
Blender scenes, Godot game prototypes. CUA/BUA for operation and verification.
Local Self-Hosting
MIT license, public on HF. Runs on SGLang / vLLM / TokenSpeed.
GLM 5.x Family Selection Guide
| Model | Input | Output | When to Use |
|---|---|---|---|
| GLM-5.3 | $1.40 | $4.40 | Hardest reasoning, maximum quality |
| GLM-5.2 | $1.40 | $4.40 | Self-hosting under MIT, stable operations |
| GLM-5.3-Flash ⬅ | $0.075 | $0.25 | Cost-sensitive, high-throughput, multimodal, coding |
💡 Strategy: Handle routine coding, agentic and document work with Flash; escalate only the hardest tasks to GLM-5.3.
Estimate Your Actual Costs
Compare token costs for GLM-5.3-Flash and all major models. Just enter your token counts.
🧮 Open Token Calculator →📊 Full Pricing Comparison · 🇨🇳 China AI Overview · 🗓️ Release Timeline · 🔗 Z.ai Official Pricing
⚠️ Disclaimer
- Pricing verified against Z.ai's official page (docs.z.ai) on 2026-08-31. Subject to change without notice. The promo price runs through 24:00 Sep 9, 2026 (UTC+8). Always check official pages before contracting.
- Benchmark figures are self-reported by Z.ai's official blog and may diverge from independent evaluations. Your mileage may vary.
- This site is informational only. We are not affiliated with or endorsed by any provider. GLM and Z.ai are trademarks of Zhipu AI.