Models
CheapestInference serves six frontier models across three pools. A subscription is per-pool: during your reserved time blocks you get unlimited usage of every model in that pool — there is no separate full-catalog tier.
- Flagship Pool — Kimi K3, Qwen3.8 Max (the two flagship-class models, from $126.65/mo with annual billing — very limited seats)
- Frontier Pool — GLM 5.2, MiniMax M3 (frontier coding & agentic models, from $50.15/mo with annual billing)
- Core Pool — DeepSeek V4 Flash, MiMo v2.5 (fast 1M-context models, from $15.29/mo with annual billing)
List models
Section titled “List models”Query the live, authoritative model list:
curl https://api.cheapestinference.com/v1/models \ -H "Authorization: Bearer YOUR_API_KEY"Each model object includes an id, owned_by, and a type field so you can filter programmatically:
{ "id": "glm-5.2", "object": "model", "created": 1677610602, "owned_by": "cheapestinference", "type": "chat"}Flagship Pool
Section titled “Flagship Pool”| Model | Model ID | Context | Per-token price elsewhere (in / out per 1M) |
|---|---|---|---|
| Kimi K3 | kimi-k3 | 256K | $3.00 / $15.00 |
| Qwen3.8 Max | qwen3.8-max | 1M | $2.00 / $6.00 |
From $126.65/mo with annual billing ($149/mo monthly). The two flagship-class models — Moonshot’s Kimi K3 and Alibaba’s Qwen3.8 Max (1M context, vision) — in one pool, very limited seats.
Frontier Pool
Section titled “Frontier Pool”| Model | Model ID | Context | Per-token price elsewhere (in / out per 1M) |
|---|---|---|---|
| GLM 5.2 | glm-5.2 | 198K | $1.40 / $4.40 |
| MiniMax M3 | minimax-m3 | 1M | $0.30 / $1.20 |
From $50.15/mo with annual billing ($59/mo monthly). The Kimi K2 family (K2.6, K2.7) previously served here has been retired; for the Kimi family, Kimi K3 is served in the Flagship Pool.
Core Pool
Section titled “Core Pool”| Model | Model ID | Context | Per-token price elsewhere (in / out per 1M) |
|---|---|---|---|
| DeepSeek V4 Flash | deepseek-v4-flash | 1M | $0.14 / $0.28 |
| MiMo v2.5 | mimo-v2.5 | ~1M | $0.14 / $0.28 |
From $15.29/mo with annual billing ($17.99–21.99/mo monthly depending on the block).
Per-token prices are the models’ list prices elsewhere — reference only, useful for comparison. On a time-block subscription you pay a flat monthly fee, not per-token charges. See Plans & Limits.
The set of models in a pool can change over time, and more pools may open.
GET /v1/modelsis always the authoritative live list.
Per-model details:
- Kimi K3 API — Moonshot’s flagship, served in the Flagship Pool
- Qwen3.8 Max API — Alibaba’s flagship with 1M context and vision, served in the Flagship Pool
- Kimi K2.7 — retired August 2026
- Kimi K2.6 — retired July 2026
- GLM 5.2 API — Zhipu’s coding & reasoning model
- MiniMax M3 API — frontier multimodal coding model with 1M context
- DeepSeek V4 Flash API — DeepSeek’s fast 1M-context model
- MiMo v2.5 API — Xiaomi’s fast, efficient ~1M-context model
- Coming — models under review — the live pipeline of models being evaluated
Using models
Section titled “Using models”Specify the model ID in your request:
# OpenAI SDKresponse = client.chat.completions.create( model="glm-5.2", # or "kimi-k3", "qwen3.8-max", "minimax-m3", "deepseek-v4-flash", ... messages=[{"role": "user", "content": "Hello"}])All models work through the OpenAI endpoint (/v1/chat/completions) and the Anthropic-compatible endpoint (/anthropic/v1/messages). The API handles format translation automatically. Your key serves the models of the pool you subscribed to.
Context windows
Section titled “Context windows”Each model has a maximum context window (shown in the pool tables above). Requests that exceed it are rejected with HTTP 400:
- Anthropic-dialect endpoints return
{"type":"error","error":{"type":"invalid_request_error","message":"prompt is too long: <tokens> tokens > <limit> maximum"}}. - OpenAI-dialect endpoints return
error.code = "context_length_exceeded".
When the model’s context window is known, responses include the x-model-context-tokens header with its value in tokens, so clients can read the exact limit for the model they are calling.
Claude Code tip: Claude Code can compact a session automatically before it fills the window. For a smooth fit, set CLAUDE_CODE_AUTO_COMPACT_WINDOW to the model’s context window — e.g. CLAUDE_CODE_AUTO_COMPACT_WINDOW=256000 for Kimi K3’s 256K window — so long sessions stay comfortably within the model’s context and keep flowing.