Skip to content

LLM Pareto Frontier — price vs intelligence, updated monthly

Living report · updated monthly · last updated
Edition: 2026-08 (latest) · all editions: 2026-08 · 2026-07

A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight models we serve and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.

Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.

$0$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Opus 4.8 GPT-5.5 Sonnet 5 GPT-5.4 Gemini 3.5 Flash Kimi K3 Qwen3.8 Max GLM 5.2 MiniMax M3 Kimi K2.6 Kimi K2.7 DeepSeek V4 Flash the top open model is now 4 pointsoff the top score — for 40% less Open weights (we serve them) Proprietary Pareto frontier

Reading left to right: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.

  • Scores across the whole board moved this edition — that’s Artificial Analysis, not the models. AA recalibrated its Intelligence Index to v4.1.1 (nine evaluations, GDPval-AA v2 and τ³-Banking among them) and every tracked model shifted up 1–5 points at unchanged prices. Compare positions, not raw scores, across editions.
  • Qwen3.8 Max enters the tracked set and the frontier in the same edition (58 at $6/M output, now open-weight and served in our Flagship Pool) — and pushes Claude Sonnet 5 (55 at $10) off the frontier. Result: Claude Opus 5 is now the only closed model left on the line.
  • Kimi K3 lands 3 index points off the top score (60 vs Claude Opus 5’s 63) — per Artificial Analysis, the narrowest open-vs-closed gap since the GLM-5 release in February — and it does it at $15/M output vs Opus 5’s $25.
  • DeepSeek V4 Flash’s new peak list pricing ($0.14/$0.28 → $0.44/$1.32) is the other mover: still frontier at 52, but the price rise puts MiniMax M3 back on the frontier as the new cheap anchor (45 at $1.20).
  • Below $10 per million output tokens, the frontier is entirely open-weight (MiniMax M3, DeepSeek V4 Flash, GLM 5.2, Qwen3.8 Max) — and every open model on the frontier, K3 included, is in our pools.
  • The near-vertical cliff at the top stays gone: the last step is Kimi K3 (60) → Opus 5 (63), 3 points for 1.7× the price. Claude Fable 5 (62) and GPT-5.6 Sol (61) remain dominated by Opus 5.
ModelAA IndexInput $/MOutput $/MWeightsPareto-efficient
Claude Opus 563$5.00$25.00closed
Claude Fable 562$10.00$50.00closeddominated by Claude Opus 5
GPT-5.6 Sol61$5.00$30.00closeddominated by Claude Opus 5
Kimi K360$3.00$15.00open
Qwen3.8 Max58$2.00$6.00open
Opus 4.857$5.00$25.00closeddominated by Claude Opus 5
GPT-5.556$5.00$30.00closeddominated by Claude Opus 5
Sonnet 555$2.00$10.00closeddominated by Qwen3.8 Max
GLM 5.253$1.40$4.40open
GPT-5.453$2.50$15.00closedmatched by GLM 5.2 at 3.4× less
DeepSeek V4 Flash52$0.44$1.32open
Gemini 3.5 Flash52$1.50$9.00closedmatched by DeepSeek V4 Flash at 6.8× less
MiniMax M345$0.30$1.20open
Kimi K2.645$0.95$4.00openmatched by MiniMax M3 at 3.3× less
Kimi K2.743$0.95$4.00opendominated by DeepSeek V4 Flash

Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.

$0.00$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index 2026-07 2026-08 (current) Current frontier Previous editions (older = fainter) Open Closed

The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.5, which we also serve, has no index entry yet and is excluded rather than estimated. Index points aren’t linear in value.

  • 2026-08 refresh (2026-08-14) — three moves in one collection. AA recalibrated its index to v4.1.1: every tracked score shifted up 1–5 points at unchanged prices — cross-edition score jumps around this date are methodology, not model changes. Qwen3.8 Max joins the tracked set (58, $2/$6) now that its weights shipped — it enters the frontier and pushes Sonnet 5 off it, leaving Opus 5 as the only closed model on the line. And DeepSeek V4 Flash’s peak list pricing took effect ($0.14/$0.28 → $0.44/$1.32), bringing MiniMax M3 back onto the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.

  • 2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.

  • 2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).

  • 2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.


All five open-weight models on the frontier — Kimi K3 and Qwen3.8 Max included — are served flat-rate in our pools — Flagship from $126.65/mo, Frontier from $50.15/mo, Core from $15.29/mo — where the marginal token costs zero during your reserved hours.