Skip to content

Qwen3.8 Max: specs, benchmarks, pricing & API options — now live in the Flagship Pool

Qwen3.8 Max (also written “Qwen 3.8 Max”; API id qwen3.8-max) is Alibaba’s flagship model, announced July 19: a 2.4-trillion-parameter system with a 1M-token context window, scoring 58 on the independent Artificial Analysis Intelligence Index (v4.1.1) — top-five territory, two points behind Kimi K3 (60), the current open-weights ceiling. As we chart below, that score at Qwen’s list price lands it on the price-vs-intelligence Pareto frontier — and knocks Claude Sonnet 5 off it.

Update, August 14, 2026: the review is over — Qwen3.8 Max is live in our Flagship Pool. Alibaba shipped the open-weight variant on August 13, the licensing gate we describe below lifted, and every Flagship subscriber can now call it with model id qwen3.8-max — unlimited, flat-rate, from $199/mo, next to Kimi K3. Setup, specs and examples: Qwen3.8 Max API — pricing, access & subscription.

Terminal window
curl https://api.cheapestinference.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3.8-max", "messages": [{"role": "user", "content": "Hello"}]}'

The analysis below is the review that got it there, kept as published (August 4) with the status lines updated.

Architecture~2.4T total parameters (MoE details unpublished)
Context window1M tokens
Open weightsReleased August 13, 2026 — Qwen3.8-2.4T-A95B, the open variant of the Max
VisionYes — image input
List price (API)$2 input / $6 output per 1M tokens (cached input from $0.25)

The benchmark picture: 58 on the Artificial Analysis index (v4.1.1) puts Qwen3.8 Max above every previous Qwen release and within two points of Kimi K3 (60). Independent benchmarking is ongoing; active-parameter counts and MoE configuration haven’t been published, so per-token compute cost can’t be derived yet.

The head-to-head everyone is asking for — the two highest-scoring models in the open(-ing) ecosystem, and on paper they’re complements rather than rivals:

Qwen3.8 MaxKimi K3
AA Intelligence Index5860
Parameters~2.4T (MoE, config unpublished)~2.8T MoE
Context window1M tokens1M tokens
List price (in / out per 1M)$2.00 / $6.00$3.00 / $15.00
Open weightsPublished August 13, 2026 (2.4T-A95B)Published July 27, 2026
Status hereLive in the Flagship Pool since August 14Live in the Flagship Pool

K3 holds the intelligence crown and the agentic-coding pedigree; Qwen3.8 Max answers with 2.5× cheaper output at two index points’ distance, plus standard sampling controls. One pool serving both covers the two profiles that matter — peak agentic reasoning and tunable long-context breadth — which is exactly what the Flagship Pool now does.

Qwen3.8 Max on the price-vs-intelligence Pareto frontier

Section titled “Qwen3.8 Max on the price-vs-intelligence Pareto frontier”

A model is on the frontier when nothing tracked is both smarter and cheaper. At index 58 for a $6.00/1M list output price, Qwen3.8 Max steps onto the frontier — five points above Claude Sonnet 5 at 40% lower list price — and pushes Sonnet 5 off it:

Served in a CheapestInference poolUnder reviewReference frontier modelsPareto frontier
4045 5055 60 $0$10 $20$30 $40$50 List output price — $ per 1M tokens AA Intelligence Index ↑ DeepSeek V4-Flash-0731 — Core Pool: index 50 at $0.28/1M — on the frontier MiniMax M3 — Frontier Pool: index 44 at $1.20/1M GLM 5.2 — Frontier Pool: index 53 at $4.40/1M — on the frontier Qwen3.8 Max — Flagship Pool: index 58 at $6.00/1M — on the frontier Gemini 3.5 Flash: index 50 at $9.00/1M Claude Sonnet 5: index 53 at $10.00/1M — pushed off the frontier by Qwen3.8 Max Kimi K3 — Flagship Pool: index 60 at $15.00/1M — on the frontier Claude Opus 4.8: index 56 at $25.00/1M Claude Opus 5: index 61 at $25.00/1M — on the frontier GPT-5.6 Sol: index 59 at $30.00/1M Claude Fable 5: index 60 at $50.00/1M (AA config: max effort, Opus 4.8 fallback) V4-Flash-0731 MiniMax M3 GLM 5.2 Gemini 3.5 Flash Sonnet 5 Qwen3.8 Max Kimi K3 Opus 4.8 Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Qwen3.8 Max: five index points above Sonnet 5 at 40% lower list output price

Four of the five models on that frontier — V4-Flash-0731, GLM 5.2, Qwen3.8 Max and Kimi K3 — are served here on flat rate. The monthly-updated, full-field version of this chart (with cost-per-task data and edition history) lives in our LLM Pareto Frontier report.

What gated the decision — and how it resolved

Section titled “What gated the decision — and how it resolved”

When this review was published (August 4), Qwen3.8 Max had no open weights and no open license. Qwen’s open releases (3.5, 3.6) shipped under Apache 2.0; its Max tier had historically stayed closed — a tension the community debated openly since the preview shipped. Our review status was explicit that licensing, not quality, was the gate.

On August 13 Alibaba resolved it: Qwen3.8-2.4T-A95B — the open-weight variant of the Max — shipped on Hugging Face and ModelScope, followed on August 14 by the dense Qwen3.8-27B under Apache 2.0. The final scorecard:

  • Quality — reviewed on real agent workloads through both our OpenAI and Anthropic endpoints, including tool calling: ✅ passed.
  • Fit — a second flagship-class model with a different profile (hybrid reasoning, standard sampling, 1M context, vision) next to Kimi K3: ✅ strong.
  • Licensing — ✅ open weights published August 13.

Result: live in the Flagship Pool on August 14 — the pipeline’s fastest gate-to-launch turnaround so far. The models-under-review board and the changelog reflect it.

Is there an unlimited Qwen3.8 Max API? Yes — live since August 14, 2026: the CheapestInference Flagship Pool serves Qwen3.8 Max with no token caps during your reserved hours, from $199/month for a daily 8-hour block, on the same subscription as Kimi K3. Model id qwen3.8-max; setup on the model page. Seats are very limited.

Does Qwen3.8 Max have open weights? Yes, as of August 13, 2026: Alibaba published Qwen3.8-2.4T-A95B, the open-weight variant of the Max, on Hugging Face and ModelScope — and the dense Qwen3.8-27B followed on August 14 under Apache 2.0.

How much does the Qwen3.8 Max API cost? Per token, list price is $2.00 per 1M input tokens and $6.00 per 1M output, with cached input from $0.25 — a 100M-token month lands around $200–600 depending on cache-hit rate and output mix. The credit-based subscription plans we analyzed in Qwen coding plans, explained cap usage per 5-hour and 7-day windows. For heavy use, our flat-rate unlimited route starts at $199/month.

What is the context window of Qwen3.8 Max? 1M tokens, per Alibaba’s published spec for the model.


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4.1 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.