⚡ Frontier intelligence at Flash cost — $0.075 in / $0.25 out (promo thru 9/9)

GLM 5.3 Flash

The first natively multimodal model in Z.ai (Zhipu)'s GLM-5 series. A 320B MoE (18B active) that outperforms GLM-5.2 on coding and agentic benchmarks at about one-tenth the price of GLM-5.3/5.2.

$0.075
Input / 1M tokens
$0.25
Output / 1M tokens
1M
Context window
128K
Max output tokens
Last updated: | Sources: Z.ai Pricing · Z.ai Model Docs · Official Blog | All Pricing · Calculator · 日本語版

Pricing (50% off promo through Sep 9, 2026)

Verified against Z.ai's official pricing page (docs.z.ai). All prices in USD per 1M tokens. Strikethrough = list price.

ModelInputCached InputCache StorageOutputNotes
GLM-5.3-Flash$0.15 $0.075$0.03 $0.015Limited-time Free$0.50 $0.2550% off. Promo ends 24:00 Sep 9, 2026 (UTC+8, Singapore time)
GLM-5.3$1.40$0.26Limited-time Free$4.40Reasoning model. ~10× Flash price
GLM-5.2$1.40$0.26Limited-time Free$4.40MIT license. Predecessor of Flash
💡 Mind the promo deadline: $0.075/$0.25 is a 50%-off promotional price ending 24:00 Sep 9, 2026 (UTC+8). After that, the currently listed rates ($0.15 input / $0.50 output) are expected to apply. Compare and estimate within the window if you're evaluating adoption.
⚠️ Thinking is always on: GLM-5.3-Flash only supports thinking.type: enabled (cannot be disabled). Thinking tokens are billed as output, so watch costs even on simple tasks. Recommended settings: reasoning_effort: max, temperature: 1, top_p: 0.95.

Primary source: docs.z.ai/guides/overview/pricing (verified 2026-08-31)

Model Specifications

Release Date
Aug 26, 2026
Model ID
glm-5.3-flash
Context Window
1M tokens
Max Output
128K tokens
Parameters
320B total / 18B active
Input Modalities
Text + Image + Video + File
Output Modality
Text
Architecture
Hybrid sparse + linear attention
License
MIT (on HuggingFace)
Provider
Zhipu AI (Z.ai)

Sources: Z.ai: GLM-5.3-Flash · Official blog · HuggingFace: zai-org/GLM-5.3-Flash

Performance — Beats GLM-5.2, approaches Claude Opus 4.8

Figures below are self-reported by Z.ai's official blog (2026-08-26) and may diverge from independent evaluations.

  • Artificial Analysis Intelligence Index v4.1.1: 57 at $0.045/task (discounted) — a level of intelligence Z.ai says previously cost ~10× more.
  • DeepSWE v1.1: 63.4 (GLM-5.2 46.2, Claude Opus 4.8 58.0, GPT-5.6 Terra 69.6, Gemini 3.7 Flash 65.3).
  • Terminal-Bench 2.1: 84.3 (close to Claude Opus 4.8 85.0, GPT-5.6 Terra 87.4, Gemini 3.7 Flash 85.8).
  • AutomationBench v1.0.6: 48.8 (GLM-5.2 26.2) — large agentic gains over GLM-5.2.
  • Z.ai Code Bench v1.0: 29.0 at max effort (nearly matching Claude Opus 4.8's 29.5).
  • vs GLM-5.2: Outperforms GLM-5.2 across all six coding/agentic benchmarks (often by a wide margin).
  • Cost-efficiency: ~1/10 the price of GLM-5.3/5.2 ($1.40/$4.40) with better performance.

Z.ai positions the model as proof that "frontier intelligence does not have to come at frontier cost." Its hybrid sparse + linear attention cuts attention compute by ~3.0× and KV cache by ~4.4× versus GLM-5.3, sharply reducing long-context serving cost. Before release it was tested anonymously as ox-alpha on OpenCode and OpenRouter, quickly becoming that week's most popular model.

Source: Z.ai blog "GLM-5.3-Flash: Frontier Intelligence, Flash Cost" (verified 2026-08-31)

Competitor Price Comparison

Comparison against low-cost Chinese models. USD per 1M tokens (standard text, no cache). GLM-5.3-Flash at promo price.

ModelProviderInputOutputvs FlashNotes
GLM-5.3-FlashZhipu (Z.ai)$0.075$0.25Baseline. Promo (thru 9/9). Multimodal
GPT-5.6 LunaOpenAI$0.20$1.20Flash cheaperFlash 62.5% cheaper input, 79% cheaper output. Luna text+image only
Gemini 3.1 Flash-LiteGoogle$0.25$1.50Flash cheaperFlash 70% cheaper input, 83% cheaper output
DeepSeek V4 FlashDeepSeek$0.22–$0.44$0.66–$1.32Flash cheaper for both input and outputPeak/off-peak. Off-peak = half of peak
MiniMax M3MiniMax$0.30$1.20Flash cheaperAfter "permanent 50% off" (3rd party)
GLM-5.3 / 5.2Zhipu (Z.ai)$1.40$4.40Flash cheaperFlash ~1/10 the price. Higher-tier reasoning models

* GLM-5.3-Flash's $0.075/$0.25 is a promo price (list $0.15/$0.50, thru 9/9). DeepSeek V4 Flash moved to peak/off-peak pricing on 2026-08-16 (cache-miss input $0.22 off-peak / $0.44 peak, output $0.66 off-peak / $1.32 peak). GPT-5.6 Luna comparison is text-input only.

Sources: official pricing pages from each provider (verified 2026-08-31). See full pricing comparison. GPT-5.6 Luna details →

Ideal Use Cases

Low cost × long context × multimodal workloads.

💻

Coding & Agents

Terminal-Bench 84.3, DeepSWE 63.4. Autonomous agents over large codebases.

🖼️

Visual Coding

Recreate UI/frontends from screenshots, recordings, and URLs with a self-verification loop.

📄

Office & Document Workflows

Generate PPTX/PDF/DOCX/XLSX. Financial research and document processing.

🔧

Tool Use & Agents

Toolathlon 78.4. Function calling and structured (JSON) output.

🎮

3D & Game Development

Blender scenes, Godot game prototypes. CUA/BUA for operation and verification.

🌐

Local Self-Hosting

MIT license, public on HF. Runs on SGLang / vLLM / TokenSpeed.

GLM 5.x Family Selection Guide

ModelInputOutputWhen to Use
GLM-5.3$1.40$4.40Hardest reasoning, maximum quality
GLM-5.2$1.40$4.40Self-hosting under MIT, stable operations
GLM-5.3-Flash ⬅$0.075$0.25Cost-sensitive, high-throughput, multimodal, coding

💡 Strategy: Handle routine coding, agentic and document work with Flash; escalate only the hardest tasks to GLM-5.3.

Estimate Your Actual Costs

Compare token costs for GLM-5.3-Flash and all major models. Just enter your token counts.

🧮 Open Token Calculator →

📊 Full Pricing Comparison · 🇨🇳 China AI Overview · 🗓️ Release Timeline · 🔗 Z.ai Official Pricing

⚠️ Disclaimer

  • Pricing verified against Z.ai's official page (docs.z.ai) on 2026-08-31. Subject to change without notice. The promo price runs through 24:00 Sep 9, 2026 (UTC+8). Always check official pages before contracting.
  • Benchmark figures are self-reported by Z.ai's official blog and may diverge from independent evaluations. Your mileage may vary.
  • This site is informational only. We are not affiliated with or endorsed by any provider. GLM and Z.ai are trademarks of Zhipu AI.