Skip to content

Qwen3.8 Max: specs, benchmarks, pricing & API options — and our review status

Qwen3.8 Max (also written “Qwen 3.8 Max”; API id qwen3.8-max) is Alibaba’s flagship model, announced July 19: a 2.4-trillion-parameter system with a 1M-token context window, scoring 53 on the independent Artificial Analysis Intelligence Index — top-five territory, four points behind Kimi K3 (57), the current open-weights ceiling. As we chart below, that score at Qwen’s list price lands it on the price-vs-intelligence Pareto frontier — and knocks Claude Sonnet 5 off it.

And to answer the question this blog exists to answer: it is officially under review for our pools. We’ve run it against real coding and agent workloads on our API stack — it behaves well — but one thing gates the decision, and it isn’t quality. More on that below, and the live status is always on our models-under-review page.

Architecture~2.4T total parameters (MoE details unpublished)
Context window1M tokens
ReasoningHybrid — reasons on demand, standard sampling params (temperature 0–2, top_p, penalties)
Open weightsAnnounced, not yet released — no license, no model card, no date
List price (API)$2 input / $6 output per 1M tokens (cached input from $0.25)

The benchmark picture: 53 on the Artificial Analysis index puts Qwen3.8 Max above every previous Qwen release and within four points of Kimi K3 (57). Independent benchmarking is ongoing; active-parameter counts and MoE configuration haven’t been published, so per-token compute cost can’t be derived yet.

The head-to-head everyone is asking for — the two highest-scoring models in the open(-ing) ecosystem, and on paper they’re complements rather than rivals:

Qwen3.8 MaxKimi K3
AA Intelligence Index5357
Parameters~2.4T (MoE, config unpublished)~2.8T MoE
Context window1M tokens1M tokens
List price (in / out per 1M)$2.00 / $6.00$3.00 / $15.00
ReasoningHybrid — toggle per requestEffort tiers low / high / max
Sampling paramsStandard (temperature 0–2, top_p, penalties)Fixed at calibrated values
Open weightsAnnounced, not yet releasedPublished July 27, 2026
Status hereUnder review — Flagship Pool candidateLive in the Flagship Pool

K3 holds the intelligence crown and the agentic-coding pedigree; Qwen3.8 Max answers with 2.5× cheaper output at two index points’ distance, plus standard sampling controls. A pool serving both would cover the two profiles that matter — peak agentic reasoning and tunable long-context breadth — which is exactly why it’s a Flagship candidate.

Qwen3.8 Max on the price-vs-intelligence Pareto frontier

Section titled “Qwen3.8 Max on the price-vs-intelligence Pareto frontier”

A model is on the frontier when nothing tracked is both smarter and cheaper. At index 53 for a $6.00/1M list output price, Qwen3.8 Max steps onto the frontier — and pushes Claude Sonnet 5 (53 at $10.00) off it:

Served in a CheapestInference poolUnder reviewReference frontier modelsPareto frontier
4045 5055 60 $0$10 $20$30 $40$50 List output price — $ per 1M tokens AA Intelligence Index ↑ DeepSeek V4-Flash-0731 — Core Pool: index 50 at $0.28/1M — on the frontier MiniMax M3 — Frontier Pool: index 44 at $1.20/1M GLM 5.2 — Frontier Pool: index 51 at $4.40/1M — on the frontier Qwen3.8 Max — under review for the Flagship Pool: index 53 at $6.00/1M — on the frontier Gemini 3.5 Flash: index 50 at $9.00/1M Claude Sonnet 5: index 53 at $10.00/1M — pushed off the frontier by Qwen3.8 Max Kimi K3 — Flagship Pool: index 57 at $15.00/1M — on the frontier Claude Opus 4.8: index 56 at $25.00/1M Claude Opus 5: index 61 at $25.00/1M — on the frontier GPT-5.6 Sol: index 59 at $30.00/1M Claude Fable 5: index 60 at $50.00/1M (AA config: max effort, Opus 4.8 fallback) V4-Flash-0731 MiniMax M3 GLM 5.2 Gemini 3.5 Flash Sonnet 5 Qwen3.8 Max Kimi K3 Opus 4.8 Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Qwen3.8 Max: Sonnet-5-class intelligence at 40% lower list output price

Three of the five models on that frontier — V4-Flash-0731, GLM 5.2 and Kimi K3 — are already served here on flat rate; Qwen3.8 Max would make it four. The monthly-updated, full-field version of this chart (with cost-per-task data and edition history) lives in our LLM Pareto Frontier report.

Qwen3.8 Max has no open weights and no open license today. Alibaba has said open weights are coming, but nothing has been published — no weights, no license text, no model card, no date. Qwen’s open releases (3.5, 3.6) shipped under Apache 2.0; its Max tier has historically stayed closed — a tension the community has been debating openly since the preview shipped. Which precedent wins here decides the outcome.

So the honest status is:

  • Quality — reviewed on real agent workloads through both our OpenAI and Anthropic endpoints, including tool calling: passes.
  • Fit — a second flagship-class model with a different profile (hybrid reasoning, standard sampling, long context) next to Kimi K3: strong.
  • Licensing — ⏳ waiting on Alibaba.

Track it live on the models-under-review page — states move from ReviewingConfirmedCapacity secured → live in the changelog. If you want your workload to count in the decision, tell us: support@cheapestinference.com.

Is there an unlimited Qwen3.8 Max API? Not yet, from anyone — including us. What exists today is per-token access at list price and quota-windowed subscription plans. Qwen3.8 Max is under review for our flat-rate unlimited pools; the live status is on the models-under-review page.

Does Qwen3.8 Max have open weights? No. As of August 2026 the model is API-only: Alibaba has announced open weights but published no license, model card, or date. Previous open Qwen releases shipped under Apache 2.0; the Max tier has historically stayed closed.

How much does the Qwen3.8 Max API cost? List price is $2.00 per 1M input tokens and $6.00 per 1M output, with cached input from $0.25. A 100M-token month lands around $200–600 depending on cache-hit rate and output mix. The credit-based subscription plans we analyzed in Qwen coding plans, explained cap usage per 5-hour and 7-day windows.

What is the context window of Qwen3.8 Max? 1M tokens, per Alibaba’s published spec for the model.


CheapestInference serves Kimi K3 (Flagship Pool), Kimi K2.7, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.