Qwen3.8 Max: specs, benchmarks, pricing & API options — and our review status
Qwen3.8 Max (also written “Qwen 3.8 Max”; API id qwen3.8-max) is Alibaba’s flagship model, announced July 19: a 2.4-trillion-parameter system with a 1M-token context window, scoring 53 on the independent Artificial Analysis Intelligence Index — top-five territory, four points behind Kimi K3 (57), the current open-weights ceiling. As we chart below, that score at Qwen’s list price lands it on the price-vs-intelligence Pareto frontier — and knocks Claude Sonnet 5 off it.
And to answer the question this blog exists to answer: it is officially under review for our pools. We’ve run it against real coding and agent workloads on our API stack — it behaves well — but one thing gates the decision, and it isn’t quality. More on that below, and the live status is always on our models-under-review page.
Qwen3.8 Max specs and benchmarks
Section titled “Qwen3.8 Max specs and benchmarks”| Architecture | ~2.4T total parameters (MoE details unpublished) |
| Context window | 1M tokens |
| Reasoning | Hybrid — reasons on demand, standard sampling params (temperature 0–2, top_p, penalties) |
| Open weights | Announced, not yet released — no license, no model card, no date |
| List price (API) | $2 input / $6 output per 1M tokens (cached input from $0.25) |
The benchmark picture: 53 on the Artificial Analysis index puts Qwen3.8 Max above every previous Qwen release and within four points of Kimi K3 (57). Independent benchmarking is ongoing; active-parameter counts and MoE configuration haven’t been published, so per-token compute cost can’t be derived yet.
Qwen3.8 Max vs Kimi K3
Section titled “Qwen3.8 Max vs Kimi K3”The head-to-head everyone is asking for — the two highest-scoring models in the open(-ing) ecosystem, and on paper they’re complements rather than rivals:
| Qwen3.8 Max | Kimi K3 | |
|---|---|---|
| AA Intelligence Index | 53 | 57 |
| Parameters | ~2.4T (MoE, config unpublished) | ~2.8T MoE |
| Context window | 1M tokens | 1M tokens |
| List price (in / out per 1M) | $2.00 / $6.00 | $3.00 / $15.00 |
| Reasoning | Hybrid — toggle per request | Effort tiers low / high / max |
| Sampling params | Standard (temperature 0–2, top_p, penalties) | Fixed at calibrated values |
| Open weights | Announced, not yet released | Published July 27, 2026 |
| Status here | Under review — Flagship Pool candidate | Live in the Flagship Pool |
K3 holds the intelligence crown and the agentic-coding pedigree; Qwen3.8 Max answers with 2.5× cheaper output at two index points’ distance, plus standard sampling controls. A pool serving both would cover the two profiles that matter — peak agentic reasoning and tunable long-context breadth — which is exactly why it’s a Flagship candidate.
Qwen3.8 Max on the price-vs-intelligence Pareto frontier
Section titled “Qwen3.8 Max on the price-vs-intelligence Pareto frontier”A model is on the frontier when nothing tracked is both smarter and cheaper. At index 53 for a $6.00/1M list output price, Qwen3.8 Max steps onto the frontier — and pushes Claude Sonnet 5 (53 at $10.00) off it:
Three of the five models on that frontier — V4-Flash-0731, GLM 5.2 and Kimi K3 — are already served here on flat rate; Qwen3.8 Max would make it four. The monthly-updated, full-field version of this chart (with cost-per-task data and edition history) lives in our LLM Pareto Frontier report.
What’s gating our decision: the weights
Section titled “What’s gating our decision: the weights”Qwen3.8 Max has no open weights and no open license today. Alibaba has said open weights are coming, but nothing has been published — no weights, no license text, no model card, no date. Qwen’s open releases (3.5, 3.6) shipped under Apache 2.0; its Max tier has historically stayed closed — a tension the community has been debating openly since the preview shipped. Which precedent wins here decides the outcome.
So the honest status is:
- Quality — reviewed on real agent workloads through both our OpenAI and Anthropic endpoints, including tool calling: passes.
- Fit — a second flagship-class model with a different profile (hybrid reasoning, standard sampling, long context) next to Kimi K3: strong.
- Licensing — ⏳ waiting on Alibaba.
Track it live on the models-under-review page — states move from Reviewing → Confirmed → Capacity secured → live in the changelog. If you want your workload to count in the decision, tell us: support@cheapestinference.com.
Common questions
Section titled “Common questions”Is there an unlimited Qwen3.8 Max API? Not yet, from anyone — including us. What exists today is per-token access at list price and quota-windowed subscription plans. Qwen3.8 Max is under review for our flat-rate unlimited pools; the live status is on the models-under-review page.
Does Qwen3.8 Max have open weights? No. As of August 2026 the model is API-only: Alibaba has announced open weights but published no license, model card, or date. Previous open Qwen releases shipped under Apache 2.0; the Max tier has historically stayed closed.
How much does the Qwen3.8 Max API cost? List price is $2.00 per 1M input tokens and $6.00 per 1M output, with cached input from $0.25. A 100M-token month lands around $200–600 depending on cache-hit rate and output mix. The credit-based subscription plans we analyzed in Qwen coding plans, explained cap usage per 5-hour and 7-day windows.
What is the context window of Qwen3.8 Max? 1M tokens, per Alibaba’s published spec for the model.
CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.