LLM Pareto Frontier — price vs intelligence, updated monthly
A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight models we serve and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.
Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.
Current frontier (edition 2026-08)
Section titled “Current frontier (edition 2026-08)”Reading left to right: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.
Analysis
Section titled “Analysis”- Scores across the whole board moved this edition — that’s Artificial Analysis, not the models. AA recalibrated its Intelligence Index to v4.1.1 (nine evaluations, GDPval-AA v2 and τ³-Banking among them) and every tracked model shifted up 1–5 points at unchanged prices. Compare positions, not raw scores, across editions.
- Qwen3.8 Max enters the tracked set and the frontier in the same edition (58 at $6/M output, now open-weight and served in our Flagship Pool) — and pushes Claude Sonnet 5 (55 at $10) off the frontier. Result: Claude Opus 5 is now the only closed model left on the line.
- Kimi K3 lands 3 index points off the top score (60 vs Claude Opus 5’s 63) — per Artificial Analysis, the narrowest open-vs-closed gap since the GLM-5 release in February — and it does it at $15/M output vs Opus 5’s $25.
- DeepSeek V4 Flash’s new peak list pricing ($0.14/$0.28 → $0.44/$1.32) is the other mover: still frontier at 52, but the price rise puts MiniMax M3 back on the frontier as the new cheap anchor (45 at $1.20).
- Below $10 per million output tokens, the frontier is entirely open-weight (MiniMax M3, DeepSeek V4 Flash, GLM 5.2, Qwen3.8 Max) — and every open model on the frontier, K3 included, is in our pools.
- The near-vertical cliff at the top stays gone: the last step is Kimi K3 (60) → Opus 5 (63), 3 points for 1.7× the price. Claude Fable 5 (62) and GPT-5.6 Sol (61) remain dominated by Opus 5.
The data
Section titled “The data”| Model | AA Index | Input $/M | Output $/M | Weights | Pareto-efficient |
|---|---|---|---|---|---|
| Claude Opus 5 | 63 | $5.00 | $25.00 | closed | ✓ |
| Claude Fable 5 | 62 | $10.00 | $50.00 | closed | dominated by Claude Opus 5 |
| GPT-5.6 Sol | 61 | $5.00 | $30.00 | closed | dominated by Claude Opus 5 |
| Kimi K3 | 60 | $3.00 | $15.00 | open | ✓ |
| Qwen3.8 Max | 58 | $2.00 | $6.00 | open | ✓ |
| Opus 4.8 | 57 | $5.00 | $25.00 | closed | dominated by Claude Opus 5 |
| GPT-5.5 | 56 | $5.00 | $30.00 | closed | dominated by Claude Opus 5 |
| Sonnet 5 | 55 | $2.00 | $10.00 | closed | dominated by Qwen3.8 Max |
| GLM 5.2 | 53 | $1.40 | $4.40 | open | ✓ |
| GPT-5.4 | 53 | $2.50 | $15.00 | closed | matched by GLM 5.2 at 3.4× less |
| DeepSeek V4 Flash | 52 | $0.44 | $1.32 | open | ✓ |
| Gemini 3.5 Flash | 52 | $1.50 | $9.00 | closed | matched by DeepSeek V4 Flash at 6.8× less |
| MiniMax M3 | 45 | $0.30 | $1.20 | open | ✓ |
| Kimi K2.6 | 45 | $0.95 | $4.00 | open | matched by MiniMax M3 at 3.3× less |
| Kimi K2.7 | 43 | $0.95 | $4.00 | open | dominated by DeepSeek V4 Flash |
How the frontier moves
Section titled “How the frontier moves”Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.
Caveats
Section titled “Caveats”The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.5, which we also serve, has no index entry yet and is excluded rather than estimated. Index points aren’t linear in value.
Changelog
Section titled “Changelog”-
2026-08 refresh (2026-08-14) — three moves in one collection. AA recalibrated its index to v4.1.1: every tracked score shifted up 1–5 points at unchanged prices — cross-edition score jumps around this date are methodology, not model changes. Qwen3.8 Max joins the tracked set (58, $2/$6) now that its weights shipped — it enters the frontier and pushes Sonnet 5 off it, leaving Opus 5 as the only closed model on the line. And DeepSeek V4 Flash’s peak list pricing took effect ($0.14/$0.28 → $0.44/$1.32), bringing MiniMax M3 back onto the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.
-
2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.
-
2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).
-
2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.
All five open-weight models on the frontier — Kimi K3 and Qwen3.8 Max included — are served flat-rate in our pools — Flagship from $126.65/mo, Frontier from $50.15/mo, Core from $15.29/mo — where the marginal token costs zero during your reserved hours.