GLM-5.3-Flash: specs, benchmarks, pricing & API options
GLM-5.3-Flash (API id glm-5.3-flash) is Z.ai’s (Zhipu AI) efficiency-tier release next to GLM-5.3, and the headline is simple: Artificial Analysis measures it at 57 on its Intelligence Index at a blended price of $0.10 per 1M tokens — $0.09 per Index task, on AA’s intelligence-vs-cost Pareto frontier, ranked #4 of 111 open-weights models it tracks. For scale: the entire index currently tops out at 63.
Unlike its big sibling — a post-training pass on the existing 743B GLM base — Flash is a new model: 320B total / 18B active MoE with hybrid sparse + linear attention, trained on a 30T-token multimodal corpus, and natively multimodal (image, video and file input) where GLM-5.3 is text-only. The weights are on Hugging Face under MIT (zai-org/GLM-5.3-Flash) — plain MIT, where the flagship’s custom license carries a security-review clause for large Model-as-a-Service operators.
GLM-5.3-Flash specs
Section titled “GLM-5.3-Flash specs”| Architecture | 320B total / 18B active MoE, hybrid sparse + linear attention |
| Multimodal | Native — image, video, text and file input; text output |
| Context window | 1M tokens; up to 128K output |
| Reasoning | Reasoning model — thinking always on, cannot be disabled on the direct API |
| Open weights | Yes — zai-org/GLM-5.3-Flash, MIT license (BF16 + FP8) |
| List price (API) | $0.15 in / $0.50 out per 1M, cached input $0.03 — 50% launch promo ($0.075 / $0.25) until September 9, 2026 |
| Measured speed | 48.6 output tok/s, 1.51s to first token (Artificial Analysis) |
The vendor benchmarks, charted
Section titled “The vendor benchmarks, charted”Z.ai’s own reported numbers (vendor harness — not independently reproduced), with the retired GLM 5.2 and the full GLM-5.3 on either side:
| Benchmark (Z.ai, vendor-run) | GLM 5.2 | GLM-5.3-Flash | GLM-5.3 |
|---|---|---|---|
| Terminal-Bench 2.1 | 81.0 | 84.3 | 88.2 |
| DeepSWE v1.1 | 46.2 | 63.4 | 66.9 |
| AutomationBench v1.0.6 | 26.2 | 48.8 | 48.2 |
Two things stand out. An 18B-active model beats the 743B GLM 5.2 on all three — a model that was frontier-class until weeks ago. And on AutomationBench it edges out the full GLM-5.3 itself. Z.ai also reports Flash within half a point of Claude Opus 4.8 on its internal Code Bench v1.0 (29.0 vs 29.5, max effort) — vendor-run, so hold it loosely until independent reproductions land.
Where it lands on the Intelligence Index
Section titled “Where it lands on the Intelligence Index”The independent signal, from Artificial Analysis (current index edition — every score below re-verified today):
Honest framing, as always: Flash does not beat the closed frontier — Claude Opus 5 leads the index at 63 — and it is three points behind the best open-weights scores (Kimi K3 and GLM-5.3, tied at 60). What is remarkable is the column AA puts next to those scores:
| Model | AA Intelligence Index | AA blended $/1M tokens |
|---|---|---|
| GLM-5.3 (max) | 60 | $0.90 |
| Qwen3.8 Max | 58 | — |
| GLM-5.3-Flash | 57 | $0.10 |
Blended = AA’s 7:2:1 cache-hit/input/output mix; scores from the current index edition. Full field and history in our monthly LLM Pareto Frontier report.
95% of GLM-5.3’s measured intelligence at a ninth of its blended price is why AA places Flash on the Pareto frontier: among everything it tracks at this intelligence level, nothing is cheaper per task. The trade-off it does not hide: speed. Flash generates ~48.6 tok/s against GLM-5.3’s ~66.6, so interactive latency is where the smaller model feels smaller.
The weights: MIT, and why that matters
Section titled “The weights: MIT, and why that matters”Z.ai shipped GLM-5.3’s weights under a custom license with a security-review clause for Model-as-a-Service operators above US$10B revenue. Flash ships under plain MIT — no clauses, BF16 and FP8 checkpoints on Hugging Face. For anyone evaluating models to serve rather than just call, that difference is not cosmetic: MIT is as serving-friendly as licenses get, and it makes Flash the most permissively-licensed near-frontier model of the moment.
Where it stands in our review
Section titled “Where it stands in our review”We serve the full GLM-5.3 unlimited in the Frontier Pool (from $71/mo, flat) since August 30. A near-frontier, MIT-licensed, 1M-context multimodal model is squarely the profile our pools exist for, and Flash is under active evaluation — the live pipeline status is always on the models-under-review page, and additions land in the changelog the day they ship.
Common questions
Section titled “Common questions”What is GLM-5.3-Flash? Z.ai’s efficiency-tier model beside GLM-5.3: a new 320B-total / 18B-active MoE with native multimodal input (image, video, file), a 1M-token context window and MIT-licensed open weights. Artificial Analysis scores it 57 on its Intelligence Index at $0.10 per 1M tokens blended.
How much does the GLM-5.3-Flash API cost? Z.ai lists $0.15 input / $0.50 output per 1M tokens (cached input $0.03), with a 50% launch discount until September 9, 2026. Artificial Analysis measures $0.09 per Intelligence-Index task and a $0.10/1M blended rate.
Does GLM-5.3-Flash have open weights? Yes — zai-org/GLM-5.3-Flash on Hugging Face under the MIT license, in BF16 and FP8. More permissive than the full GLM-5.3, whose custom license adds a security-review clause for large Model-as-a-Service operators.
How good is GLM-5.3-Flash at coding? On Z.ai’s vendor benchmarks it scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1 — above the retired 743B GLM 5.2 on both — and 48.8 on AutomationBench, marginally above the full GLM-5.3. Independently, its 57 on the AA index is three points behind the best open-weights models. Its measured output speed is ~48.6 tok/s.
Is there an unlimited GLM-5.3-Flash API subscription? Not from us today — Flash is under evaluation (live status). The full GLM-5.3 is served unlimited in the Frontier Pool from $71/mo: flat monthly fee, no token caps during your reserved hours.
CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.