Skip to content

GLM-5.3: specs, benchmarks, pricing & API options — and our review status

GLM-5.3 (API id glm-5.3) is Z.ai’s (Zhipu AI) new coding and agentic model, released today, August 14, 2026, under the tagline “Built to Code. Ready for Cyber Defense.” The architecture story is unusual and worth being precise about: GLM-5.3 keeps the same 743B-parameter base model as GLM 5.2 — every reported gain comes from scaled-up post-training alone. On Z.ai’s own benchmark suite that post-training buys a lot: it calls GLM-5.3 the strongest open-weights coding model it has measured, and reports a cyber-security capability that grew faster than the company anticipated.

And the question this blog exists to answer: GLM-5.3 is officially under review for our pools, as of today. We already serve GLM 5.2 in the Frontier Pool, so 5.3 enters the pipeline as the natural upgrade candidate for that slot. What gates the decision is not quality signals — it’s that the model is API-only today: open weights are promised roughly two weeks out, after Z.ai completes its own safety evaluation. Live status is always on our models-under-review page.

Base modelSame 743B base as GLM 5.2 — not a new pretrain; gains from extended post-training
Context windowZ.ai advertises a 1M-token variant (glm-5.3[1m], with context compaction); the standard-API spec is not yet published
ReasoningEffort levels low / high / max — default max; thinking cannot be disabled on the direct API
Open weightsPromised ~2 weeks after launch, pending Z.ai’s safety evaluation — no weights, no license, no model card today
LicenseUnpublished. GLM 5.2 shipped MIT; that precedent does not automatically set 5.3’s terms
List price (API)Not yet published — GLM 5.2 lists $1.40 in / $4.40 out per 1M as reference
Availability todayFirst-party API access from Z.ai only — no open weights, no third-party serving yet; works with Claude Code, OpenCode, and Codex via compatible endpoints

GLM-5.3 benchmarks: what post-training bought

Section titled “GLM-5.3 benchmarks: what post-training bought”

All numbers below are Z.ai’s own reported results — vendor-run, not yet independently reproduced, and the Artificial Analysis index hasn’t rated GLM-5.3 yet. With that caveat on the table, the GLM 5.2 → GLM-5.3 deltas are the story, because the base model is identical:

BenchmarkGLM 5.2GLM-5.3Δ
Terminal-Bench 2.181.088.2+9%
Terminal-Bench 3.04.628.3+515%
DeepSWE v1.146.266.9+45%
SWE-Marathon v1.119.442.5+119%
FrontierSWE67.578.1+16%
NL2Repo48.958.0+19%
Toolathlon Verified59.973.0+22%
AutomationBench v1.0.626.248.2+84%
CyberGym77.284.5+9%

Two readings. The charitable one: the biggest jumps land on the newest, hardest agentic benchmarks (Terminal-Bench 3.0, SWE-Marathon) — exactly where post-training on agent trajectories should show up, and exactly the workloads coding agents run all day. The skeptical one: several of these benchmarks are new or Z.ai-adjacent, and until independent runs land, “strongest open-weights coding model” is a claim, not a fact. Both readings can wait two weeks — the open-weights release is when independent verification becomes possible.

The cyber-defense angle — and why the weights are two weeks out

Section titled “The cyber-defense angle — and why the weights are two weeks out”

The unusual part of this launch is that Z.ai leads with cyber security as a first-class capability, not a footnote. It reports GLM-5.3 at 84.5 on CyberGym — above its figures for Claude Mythos 5 (83.8) and GPT-5.6 Sol (83.6) — and says the model found thousands of real vulnerabilities across open-source projects during training. Z.ai’s framing is defensive: vulnerability detection at scale.

That capability is also the stated reason the weights aren’t out yet. Rather than shipping weights on day one — as it did with GLM 5.2 — Z.ai is running a staged release: API first, then open weights roughly two weeks after launch, once its own safety evaluation and hardening work is complete. Whatever you think of the trade-off, it’s a more deliberate open-weights process than the ecosystem norm, and it puts a concrete clock on the one thing our review is waiting for.

GLM-5.3 vs GLM 5.2 — the model we serve today

Section titled “GLM-5.3 vs GLM 5.2 — the model we serve today”
GLM-5.3GLM 5.2
Base743B (same base)743B
What’s newScaled post-training: agentic coding, tool use, cyber
Context1M via Z.ai’s [1m] variant; standard API unpublished198K as served here
ReasoningEffort low / high / max, thinking always onStandard GLM 5.2 semantics
Open weightsPromised ~2 weeks post-launchPublished (MIT)
List price (per 1M)Not yet published$1.40 in / $4.40 out
Status hereUnder review — Frontier candidateLive in the Frontier Pool

Because the base is unchanged, this isn’t a “new model vs old model” decision so much as a post-training upgrade — the same shape as DeepSeek’s V4-Flash-0731 build, which we upgraded in place in the Core Pool within days of release. If GLM-5.3’s weights land with a usable license and it passes our quality evaluation on real coding and agent workloads, the natural outcome is the same: the Frontier Pool’s GLM slot upgrades, and every existing subscription simply gets the better model.

So the honest status board:

  • Quality — vendor numbers are strong; our own evaluation on real agent workloads (both OpenAI and Anthropic endpoints, tool calling included) runs in parallel: ⏳ in progress.
  • Fit — a post-training upgrade of a model already serving Frontier Pool workloads: as clean as fit gets.
  • Licensing — ⏳ waiting on Z.ai’s open-weights release and license text, on the ~two-week clock Z.ai itself set.

Track it live on the models-under-review page — states move Reviewing → Confirmed → Capacity secured → live in the changelog. If you want your workload to count in the decision: support@cheapestinference.com.

Is there an unlimited GLM-5.3 API? Not yet, from anyone. GLM-5.3 is API-only and first-party today. It is under review for our flat-rate unlimited pools as a Frontier Pool candidate; meanwhile, GLM 5.2 — same base model — is served unlimited on time-block subscriptions today.

Does GLM-5.3 have open weights? Not yet. Z.ai has committed to releasing open weights roughly two weeks after the August 14 launch, once its internal safety evaluation of the model’s cyber capabilities is complete. No license text or model card has been published; GLM 5.2’s MIT license doesn’t automatically carry over.

How much does the GLM-5.3 API cost? Per-token list pricing hasn’t been published — the closest reference is GLM 5.2’s list rate ($1.40 in / $4.40 out per 1M) until Z.ai posts 5.3 rates.

What is the difference between GLM-5.3 and GLM 5.2? Same 743B base model — GLM-5.3 is extended post-training on top of it, targeting agentic coding, tool use, and cyber-security workloads. Z.ai reports large gains on agentic benchmarks (SWE-Marathon 19.4 → 42.5, Terminal-Bench 3.0 4.6 → 28.3); all numbers are vendor-run so far.

What is the context window of GLM-5.3? Z.ai advertises 1M tokens via the glm-5.3[1m] variant with context compaction. The standard-API context spec hasn’t been published yet.


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.2 and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.