LLM API Pricing Comparison — per-token prices over time, updated monthly
What a million tokens costs, tracked over time. Two data sources, both dated and verifiable: historical months reconstructed from Wayback Machine snapshots of Artificial Analysis model pages (per-model snapshot timestamps preserved in the underlying data), and the same live monthly collection that feeds the Pareto Frontier report from July 2026 on. Where no snapshot exists, the chart shows a gap — never an estimate; the archive’s coverage of AA starts February 2026. Superseded models (dashed) stay on the chart, because model churn is half the price story.
The numbers
Section titled “The numbers”Sorted by the metric that actually matters — dollars of output per Intelligence Index point (lower = more intelligence per dollar):
| Model | Output $/M | Input $/M | AA Index | $ per index point | Weights |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | $1.32 | $0.44 | 52 | $0.025 | open |
| MiniMax M3 | $1.20 | $0.30 | 45 | $0.027 | open |
| GLM 5.2 | $4.40 | $1.40 | 53 | $0.083 | open |
| Kimi K2.6 | $4.00 | $0.95 | 45 | $0.089 | open |
| Kimi K2.7 | $4.00 | $0.95 | 43 | $0.093 | open |
| Qwen3.8 Max | $6.00 | $2.00 | 58 | $0.103 | open |
| Gemini 3.5 Flash | $9.00 | $1.50 | 52 | $0.173 | closed |
| Sonnet 5 | $10.00 | $2.00 | 55 | $0.182 | closed |
| Kimi K3 | $15.00 | $3.00 | 60 | $0.250 | open |
| GPT-5.4 | $15.00 | $2.50 | 53 | $0.283 | closed |
| Claude Opus 5 | $25.00 | $5.00 | 63 | $0.397 | closed |
| Opus 4.8 | $25.00 | $5.00 | 57 | $0.439 | closed |
| GPT-5.6 Sol | $30.00 | $5.00 | 61 | $0.492 | closed |
| GPT-5.5 | $30.00 | $5.00 | 56 | $0.536 | closed |
| Claude Fable 5 | $50.00 | $10.00 | 62 | $0.806 | closed |
| MiMo V2.5 | $0.28 | $0.14 | — | — | open |
Analysis
Section titled “Analysis”- Below the flagship tier, every open model still buys intelligence cheaper than every closed model. The most efficient closed option (Gemini 3.5 Flash, $0.17 per index point) costs about twice the least efficient of the value-tier open models (Kimi K2.7, $0.093) — and 7× DeepSeek V4 Flash’s $0.025, a gap DeepSeek’s own peak-pricing move narrowed from 26× last collection.
- Kimi K3 broke the pattern — the first open model priced like a closed one. At $15/M output ($0.250 per index point) it sits between Sonnet 5 and GPT-5.4 on the board. What it buys: an index of 60, above everything closed except Opus 5, Fable 5 and GPT-5.6 Sol. Qwen3.8 Max undercuts that trade at $6/M for an index of 58.
- The premium is still at the top. From GLM 5.2 to Claude Opus 5 the index rises 19% (53 → 63) while the price per point rises 4.8× ($0.083 → $0.397) — and Fable 5 sits at $0.806/point. You don’t pay for intelligence — you pay for the last few points of it.
- What to watch as editions accumulate: open-weight prices trend down because any host can serve the weights and competition compresses margins to serving cost; closed prices only move when the sole vendor decides. A year ago the best open model scored in the low 20s on this index; today it’s 60 — though with K3, for the first time, the top open price moved up too, and DeepSeek’s peak pricing shows even the floor can rise. The value tier (GLM 5.2 and below) keeps getting smarter without getting much pricier.
Break-even: when does flat-rate beat the meter?
Section titled “Break-even: when does flat-rate beat the meter?”The per-token prices above stop mattering past a usage threshold — here it is, computed from the table (blended at the 80% input / 20% output mix typical of real workloads):
| Flat subscription | vs paying per token for | Blended $/M | Break-even: tokens/month | In agent-hours¹ |
|---|---|---|---|---|
| Frontier, from $50.15/mo | GLM 5.2 itself (open, list price) | $2.00 | 24.2M | ~12 h/month |
| Frontier, from $50.15/mo | GPT-5.4 — same index, closed | $5.00 | 9.7M | ~5 h/month |
| Frontier, from $50.15/mo | GPT-5.5 | $10.00 | 4.8M | ~2 h/month |
| Core, from $15.29/mo | DeepSeek V4 Flash itself (open, peak list price) | $0.62 | 24.7M | ~12 h/month |
| Core, from $15.29/mo | GPT-5.4 mini — same index, closed | $1.50 | 10.2M | ~5 h/month |
¹ We measured a coding agent at ~2.1M tokens/hour — so a single working day of agent traffic per month clears every break-even on this table against closed per-token pricing. Against the open models’ own list prices the bar is higher: about a day and a half for Frontier vs GLM 5.2, and — since DeepSeek’s peak repricing — about the same for Core vs DeepSeek’s peak rate. Below those volumes, the meter is genuinely cheaper; that’s the honest threshold.
The same thing as a picture — monthly cost against monthly usage; where a per-token diagonal crosses the flat line, the subscription starts winning:
One structural note: the flat side of this comparison is a reserved daily time block serving one request at a time per key — you’re buying capacity, not metered volume, so past break-even the marginal token costs zero for the rest of the month. The full framework for that decision is your AI bill should scale with users, not usage.
Caveats
Section titled “Caveats”Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight models are often cheaper on aggregators, so the open rows are conservative. ”$ per index point” divides output price by a composite index — it’s a comparison heuristic, not a claim that index points are linear in value. MiMo V2.5, which we serve, appears with its current OpenRouter price but no index or history — it has no Artificial Analysis entry yet, so those cells stay empty rather than estimated.
Changelog
Section titled “Changelog”-
2026-08-27 — Core pool block reprice: blocks now $17.99–$21.99/mo (cheapest block was $16.49); annual anchor $14.02 → $15.29/mo. Break-even rows recomputed on the new anchor.
-
2026-08-14 — DeepSeek V4 Flash $0.14/$0.28 → $0.44/$1.32: the peak/off-peak pricing DeepSeek announced for August 16 is now what Artificial Analysis lists (peak rate shown; off-peak stays near the old level). Qwen3.8 Max joins the board (58, $2/$6 — $0.103/point) now that its weights are open. All AA index scores shifted up 1–5 points with AA’s v4.1.1 recalibration; the $-per-point column is recomputed on the new scores.
-
2026-08-14 — Core pool block prices updated ($16.49–$21.99/mo per block; annual anchor $12.74 → $14.02/mo); break-even rows recomputed. DeepSeek announced peak/off-peak API pricing effective Aug 16 (V4 Flash output $0.28 → $0.66–$1.32/M) — the per-token board picks it up in the next monthly collection.
-
2026-07-28 — release-week collection: Kimi K3 ($3/$15), Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) join the board. Sonnet 5’s price was cut $3/$15 → $2/$10 — the first closed price drop on this tracker’s watch. K3 is the first open model priced in closed mid-tier territory.
-
2026-07-22 — break-even table and chart recomputed on the pool anchors current at the time (Frontier from $50.15/mo, Core at $12.74/mo with annual billing — since raised to $14.02/mo, see the 2026-08-14 entry); MiMo V2.5’s OpenRouter input price rose $0.105 → $0.14/M (output unchanged at $0.28/M).
-
2026-07-06 (backfill) — historical prices February–June 2026 reconstructed from Wayback Machine snapshots of Artificial Analysis: 19 models including superseded ones (Opus 4.6→4.7 handover, Kimi K2.5, GLM 5.1, MiniMax M2, Gemini 3 Pro/Flash). Notable finds in the record: DeepSeek V3.2’s output price rose from $0.42 to $1.60/M in May — open-weight prices mostly fall, but not always — and Kimi K2.5 wobbled $3.00 → $2.85 → $3.00.
-
2026-07-06 (first edition) — baseline prices ingested from the 2026-07 Pareto collection. Range on the board: $0.28/M (DeepSeek V4 Flash) to $50/M (Claude Fable 5) per million output tokens — a 178× spread.
On our pools these per-token prices stop applying at all: flat monthly, no token caps during your reserved hours — Flagship from $126.65/mo, Frontier from $50.15/mo, Core from $15.29/mo. See the pools or read why your AI bill should scale with users, not usage.