Kimi K3: specs, benchmarks, and our day-one plan to serve it
Kimi K3 is Moonshot AI’s new flagship, announced on July 16 after a leaked promotion page on Moonshot’s own platform tipped the release a day early. The headline is simple: on the independent Artificial Analysis Intelligence Index it scores 57 — above Claude Opus 4.8 (56). If the promised open-weight release lands on schedule, K3 becomes the first open-weight model to outscore Anthropic’s Opus tier on that index.
And to answer the question this blog exists to answer: yes, we will serve it. We have already secured and confirmed the capacity to run a model of this size. K3 goes live on CheapestInference the moment two things are true: the weights are actually published, and the license terms permit commercial serving. Nothing else is in the way.
What Moonshot announced
Section titled “What Moonshot announced”| Architecture | Mixture-of-Experts, ~2.8T total parameters |
| Context window | 1M tokens |
| Input | Text, image, and video |
| Variants at launch | K3 Max (chat and agent tasks) · K3 Swarm Max (large-scale parallel processing) |
| Available today | Moonshot’s API, Kimi Code, and the Kimi app |
| Open weights | Promised by July 27, 2026 |
| List price (API) | $3 input / $15 output per 1M tokens |
Coverage from the launch day: TechCrunch on the Opus gap closing, Fortune on Chinese AI entering Fable-level territory, and Simon Willison’s notes for a practitioner’s first look.
The numbers so far
Section titled “The numbers so far”- Artificial Analysis Intelligence Index: 57. For scale: Claude Opus 4.8 scores 56, and the best open-weight model until now — GLM 5.2 — scores 51. If the weights ship, the open ceiling jumps six points in one release. Where every model sits on price-vs-intelligence is our Pareto Frontier report, and the August edition will chart K3 the moment it qualifies.
- #1 on Frontend Code Arena with 1,679 points — ahead of Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and GLM 5.2 (1,587).
- The Kimi track record. Moonshot’s K2.6 still holds the best open SWE-bench Verified score (80.2), and the K2 line has been the default open choice for tool-heavy agent work — see the tier-fair matchups in our Which-LLM guide.
The honest caveats, same as in our living reports: the index is one composite and task-specific rankings differ; these are launch-week numbers, mostly on Moonshot’s own serving stack; and the weights are not downloadable yet — the K2 line shipped under a Modified MIT license, but K3’s exact license text isn’t published. Our State of Open Weights roster only updates when a model is real, downloadable, and served — K3 gets its row when it clears that bar.
What it changes if it lands
Section titled “What it changes if it lands”Today, every model above GLM 5.2’s intelligence score is closed, and every proprietary model on the Pareto frontier is Anthropic’s. An open-weight model at 57 — with a $3/$15 list price undercutting every closed model in its class — pushes the open frontier right up against the closed tier for the first time. That’s not an incremental release; it’s the open-vs-closed gap, which our reports have tracked at 5+ index points all year, compressing to one.
Our plan: day one, pricing announced at launch
Section titled “Our plan: day one, pricing announced at launch”- Capacity: secured. The infrastructure to serve K3 is confirmed and waiting. When Moonshot publishes the weights and the license clears commercial serving, we ship — we’re not starting the clock that day.
- Pricing: deliberately open until launch. Depending on how K3 performs on real workloads, it may join one of the existing pools — or debut in a dedicated pool with pricing exclusive to this model. We’ll make that call on real numbers, not launch-week hype, and announce it on the pools page, which is always the live source for lineup and prices.
- The usual, once live: one OpenAI- and Anthropic-compatible API, flat-rate time-block subscriptions, no token caps during your reserved hours — so it drops into Claude Code, Cline, or any compatible client the day it appears in
GET /v1/models.
If you want to be running K3 the week it exists, create an account now — the model list updates live, no waitlist.
CheapestInference serves Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (Frontier Pool, from $48.45/mo billed annually) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool, from $12.74/mo) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.