Skip to content

State of Open Weights — the open models we serve, and why

Living report · updated monthly (numbers) · on lineup changes (roster) · last updated

Open a model router today and you’ll find 400+ models — the same family in five quantizations, three deprecation states, a dozen providers with different latencies and silent fallbacks. Choice isn’t the product; it’s a tax. This report is the opposite: the small set of open-weight models we actually operate, each one there for a job, with the facts that matter — real license terms, architecture, context, current prices and scores — kept current so you never have to reconcile stale numbers across blog posts.

Sources: model facts from the official Hugging Face cards (linked per row); prices and Intelligence Index from Artificial Analysis, refreshed monthly by the same collector that updates the Pareto Frontier report.

ModelPoolThe jobParamsContextLicense
Kimi K3FlagshipOpen flagship — top open-weights score on the AA index (60, v4.1.1)2.8T MoE / 104B active1MKimi K3 License³
Kimi K2.7— (retired August 2026)Agentic coding, tool use at a fraction of K3’s price1T MoE / 32B active256KModified MIT¹
Kimi K2.6— (retired July 2026)Best open SWE-bench Verified score; pinned agent workflows1T MoE / 32B active256KModified MIT¹
GLM 5.2FrontierLong-horizon coding; best open intelligence-per-dollar on the frontier753B MoE1MMIT
MiniMax M3Frontier1M context + native multimodal input~428B MoE / 23B active1MCommunity License²
DeepSeek V4 FlashCoreHigh-volume workloads: extraction, chat, summarization, RAG284B MoE / 13B active1MMIT
MiMo V2.5CoreBudget multimodal; strongest coding card in its price class310B MoE / 15B active1MMIT

¹ Modified MIT: attribution required only above 100M MAU or $20M/mo revenue. ² MiniMax Community License: “Built with MiniMax M3” attribution; >$20M/yr revenue requires written authorization. ³ Kimi K3 License: Moonshot’s own license for the K3 weights — read the license text before self-hosting. “Open” is not one thing — MIT and Apache 2.0 mean unrestricted commercial use; community licenses sit between open and proprietary. For API consumers none of this matters day to day; it matters if you later self-host or embed weights in a product.

ModelInput $/MOutput $/MAA IndexHeadline benchmark
Kimi K3$3$1560Terminal-Bench 2.1 88.3 · GPQA Diamond 93.5
Qwen3.8 Max$2$658SWE-bench Pro 67.7 · Terminal-Bench 2.1 86.6
Kimi K2.7$0.95$443MCP Mark Verified 81.1 (vendor suite)
Kimi K2.6$0.95$445SWE-bench Verified 80.2 · Pro 58.6
GLM 5.2$1.4$4.453SWE-bench Pro 62.1 · Terminal-Bench 2.1 81.0
MiniMax M3$0.3$1.245SWE-bench Verified 80.5 · Pro 59.0
DeepSeek V4 Flash$0.44$1.3252Terminal-Bench 2.1: 82.7 (0731, vendor)
MiMo V2.5$0.14$0.28SWE-bench Pro 56.1 · Terminal-Bench 2 65.8

Prices are the vendors’ list prices per million tokens as tracked by Artificial Analysis (MiMo V2.5: OpenRouter listing — no AA entry yet). Benchmark scores are as published on each model’s official card; scaffolding differs between labs. Where these models sit against the closed flagships — and which are Pareto-efficient — is the Pareto Frontier report; tier-fair matchups are in the Which-LLM guide.

Flagship when you want the strongest open models there are — Kimi K3 and Qwen3.8 Max, 3 and 5 points off the top closed score. Frontier when the work is agentic coding, multi-step agents, or anything where a failed run costs more than the tokens did. Core when the work is volume — extraction, classification, summarization, chat, pipelines that run all day. All three are flat-rate: no token caps during your reserved hours, so the per-token prices above stop mattering once you’re inside your window.

  • 2026-08-14Qwen3.8 Max joins the roster in the Flagship Pool: its open weights (Qwen3.8-2.4T-A95B, 2.4T MoE / 95B active) shipped August 12 under the custom Qwen3.8-Max License, lifting the licensing gate that held it in review. Kimi K2.7 retired from the Frontier Pool in the same window. Board-wide: AA’s Intelligence Index recalibration to v4.1.1 lifted every score 1–5 points at unchanged prices (K3 60, Qwen 58, GLM 5.2 53), and DeepSeek V4 Flash’s list price moved to its new peak rate ($0.44/$1.32).

  • 2026-08-01 — DeepSeek V4 Flash re-scored 40 → 50 on the AA index at unchanged prices, following the V4-Flash-0731 retrain.

  • 2026-07-28Kimi K3 joins the roster in the new Flagship Pool: 2.8T MoE (104B active), 1M context, index 57 — 4 points off the top closed model, the narrowest open-vs-closed gap since February per Artificial Analysis. Kimi K2.6’s index nudged 43 → 44 in the same collection.

  • 2026-07-27 — Kimi K2.6 retired from the Frontier Pool; Kimi K2.7 remains its successor in the same pool.

  • 2026-07-22 — MiMo V2.5 input price on its OpenRouter listing rose $0.105 → $0.14/M; output unchanged at $0.28/M. Still no Artificial Analysis entry.

  • 2026-07-06 (first edition) — baseline roster: Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3 (Frontier); DeepSeek V4 Flash, MiMo V2.5 (Core). This report absorbs the per-model tables previously maintained in blog posts, which now link here instead of carrying their own copies of the numbers.


The live pool menu — with current block prices and availability — is always at cheapestinference.com/pools: Flagship from $126.65/mo, Frontier from $50.15/mo, Core from $15.29/mo, one OpenAI- and Anthropic-compatible API.