Skip to content

Kimi K3: specs, benchmarks, and our day-one plan to serve it

Kimi K3 is Moonshot AI’s new flagship, announced on July 16 after a leaked promotion page on Moonshot’s own platform tipped the release a day early. The headline is simple: on the independent Artificial Analysis Intelligence Index it scores 57 — above Claude Opus 4.8 (56). If the promised open-weight release lands on schedule, K3 becomes the first open-weight model to outscore Anthropic’s Opus tier on that index.

And to answer the question this blog exists to answer: yes, we will serve it. We have already secured and confirmed the capacity to run a model of this size. K3 goes live on CheapestInference the moment two things are true: the weights are actually published, and the license terms permit commercial serving. Nothing else is in the way.

ArchitectureMixture-of-Experts, ~2.8T total parameters
Context window1M tokens
InputText, image, and video
Variants at launchK3 Max (chat and agent tasks) · K3 Swarm Max (large-scale parallel processing)
Available todayMoonshot’s API, Kimi Code, and the Kimi app
Open weightsPromised by July 27, 2026
List price (API)$3 input / $15 output per 1M tokens

Coverage from the launch day: TechCrunch on the Opus gap closing, Fortune on Chinese AI entering Fable-level territory, and Simon Willison’s notes for a practitioner’s first look.

  • Artificial Analysis Intelligence Index: 57. For scale: Claude Opus 4.8 scores 56, and the best open-weight model until now — GLM 5.2 — scores 51. If the weights ship, the open ceiling jumps six points in one release. Where every model sits on price-vs-intelligence is our Pareto Frontier report, and the August edition will chart K3 the moment it qualifies.
  • #1 on Frontend Code Arena with 1,679 points — ahead of Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and GLM 5.2 (1,587).
  • The Kimi track record. Moonshot’s K2.6 still holds the best open SWE-bench Verified score (80.2), and the K2 line has been the default open choice for tool-heavy agent work — see the tier-fair matchups in our Which-LLM guide.

The honest caveats, same as in our living reports: the index is one composite and task-specific rankings differ; these are launch-week numbers, mostly on Moonshot’s own serving stack; and the weights are not downloadable yet — the K2 line shipped under a Modified MIT license, but K3’s exact license text isn’t published. Our State of Open Weights roster only updates when a model is real, downloadable, and served — K3 gets its row when it clears that bar.

Today, every model above GLM 5.2’s intelligence score is closed, and every proprietary model on the Pareto frontier is Anthropic’s. An open-weight model at 57 — with a $3/$15 list price undercutting every closed model in its class — pushes the open frontier right up against the closed tier for the first time. That’s not an incremental release; it’s the open-vs-closed gap, which our reports have tracked at 5+ index points all year, compressing to one.

Our plan: day one, pricing announced at launch

Section titled “Our plan: day one, pricing announced at launch”
  • Capacity: secured. The infrastructure to serve K3 is confirmed and waiting. When Moonshot publishes the weights and the license clears commercial serving, we ship — we’re not starting the clock that day.
  • Pricing: deliberately open until launch. Depending on how K3 performs on real workloads, it may join one of the existing pools — or debut in a dedicated pool with pricing exclusive to this model. We’ll make that call on real numbers, not launch-week hype, and announce it on the pools page, which is always the live source for lineup and prices.
  • The usual, once live: one OpenAI- and Anthropic-compatible API, flat-rate time-block subscriptions, no token caps during your reserved hours — so it drops into Claude Code, Cline, or any compatible client the day it appears in GET /v1/models.

If you want to be running K3 the week it exists, create an account now — the model list updates live, no waitlist.


CheapestInference serves Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (Frontier Pool, from $48.45/mo billed annually) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool, from $12.74/mo) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.