Skip to content

Models

CheapestInference serves six frontier models across three pools. A subscription is per-pool: during your reserved time blocks you get unlimited usage of every model in that pool — there is no separate full-catalog tier.

  • Flagship Pool — Kimi K3, Qwen3.8 Max (the two flagship-class models, from $126.65/mo with annual billing — very limited seats)
  • Frontier Pool — GLM 5.2, MiniMax M3 (frontier coding & agentic models, from $50.15/mo with annual billing)
  • Core Pool — DeepSeek V4 Flash, MiMo v2.5 (fast 1M-context models, from $15.29/mo with annual billing)

Query the live, authoritative model list:

Terminal window
curl https://api.cheapestinference.com/v1/models \
-H "Authorization: Bearer YOUR_API_KEY"

Each model object includes an id, owned_by, and a type field so you can filter programmatically:

{
"id": "glm-5.2",
"object": "model",
"created": 1677610602,
"owned_by": "cheapestinference",
"type": "chat"
}
ModelModel IDContextPer-token price elsewhere (in / out per 1M)
Kimi K3kimi-k3256K$3.00 / $15.00
Qwen3.8 Maxqwen3.8-max1M$2.00 / $6.00

From $126.65/mo with annual billing ($149/mo monthly). The two flagship-class models — Moonshot’s Kimi K3 and Alibaba’s Qwen3.8 Max (1M context, vision) — in one pool, very limited seats.

ModelModel IDContextPer-token price elsewhere (in / out per 1M)
GLM 5.2glm-5.2198K$1.40 / $4.40
MiniMax M3minimax-m31M$0.30 / $1.20

From $50.15/mo with annual billing ($59/mo monthly). The Kimi K2 family (K2.6, K2.7) previously served here has been retired; for the Kimi family, Kimi K3 is served in the Flagship Pool.

ModelModel IDContextPer-token price elsewhere (in / out per 1M)
DeepSeek V4 Flashdeepseek-v4-flash1M$0.14 / $0.28
MiMo v2.5mimo-v2.5~1M$0.14 / $0.28

From $15.29/mo with annual billing ($17.99–21.99/mo monthly depending on the block).

Per-token prices are the models’ list prices elsewhere — reference only, useful for comparison. On a time-block subscription you pay a flat monthly fee, not per-token charges. See Plans & Limits.

The set of models in a pool can change over time, and more pools may open. GET /v1/models is always the authoritative live list.

Per-model details:

Specify the model ID in your request:

# OpenAI SDK
response = client.chat.completions.create(
model="glm-5.2", # or "kimi-k3", "qwen3.8-max", "minimax-m3", "deepseek-v4-flash", ...
messages=[{"role": "user", "content": "Hello"}]
)

All models work through the OpenAI endpoint (/v1/chat/completions) and the Anthropic-compatible endpoint (/anthropic/v1/messages). The API handles format translation automatically. Your key serves the models of the pool you subscribed to.

Each model has a maximum context window (shown in the pool tables above). Requests that exceed it are rejected with HTTP 400:

  • Anthropic-dialect endpoints return {"type":"error","error":{"type":"invalid_request_error","message":"prompt is too long: <tokens> tokens > <limit> maximum"}}.
  • OpenAI-dialect endpoints return error.code = "context_length_exceeded".

When the model’s context window is known, responses include the x-model-context-tokens header with its value in tokens, so clients can read the exact limit for the model they are calling.

Claude Code tip: Claude Code can compact a session automatically before it fills the window. For a smooth fit, set CLAUDE_CODE_AUTO_COMPACT_WINDOW to the model’s context window — e.g. CLAUDE_CODE_AUTO_COMPACT_WINDOW=256000 for Kimi K3’s 256K window — so long sessions stay comfortably within the model’s context and keep flowing.