Skip to content

Flat Rate vs Per-Token — the break-even curve and the savings factor, updated monthly

Living report · updated monthly, with each Pareto edition, and on every price change · last updated

Every conversation about inference cost lands on the same question: at what point does a flat monthly block beat paying per token? The answer is a curve, not a number. This report draws it, computes the break-even volume for every pool from the vendors’ own published per-token list prices, and states the savings factor honestly — including the ceiling a single key will not go past.

Two inputs, both live: our flat prices (Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $18.70/mo, the cheapest block of each pool billed annually) and the per-token list prices tracked in the LLM Price Tracker. Every figure below is recomputed at build time from those sources; nothing is hand-typed.

Per-token cost is a straight line through the origin: twice the tokens, twice the bill. A flat block is a horizontal line: the invoice is decided when you subscribe. The two cross at the break-even volume. To the left, per-token is cheaper. To the right, the flat block is, and the gap widens with every extra token.

Per-token, list price of the best comparable alternativeFlat Core block, from $18.70/mo
0100M200M300M400M500M tokens per month $0$50$100 Per-token: $0.17 per million tokens, blended 80% input / 20% output Flat Core block: $18.70 per month, unlimited tokens during the reserved hours Break-even: about 111M tokens per month, roughly 3.7M per day break-even ≈ 111M tokens/month per-token: ≈ $75.60 at 450M flat block: $18.70, any volume

The chart shows the Core pool; the other pools follow the same shape at their own prices (table below).

Break-even per model

One row per pool, against the published per-token list price of the best comparable alternative — the conservative case. Blended rate: 80% input and 20% output, no cache discount. Flat price is the cheapest block billed annually; the last column is the same break-even billed monthly. Prices as of edition 2026-09; current pricing is on the pools page.

PoolBlended per-token list priceBreak-even volumePer day (30-day month)Billed monthly
Flagship, from $169.15/mo ≈ $2.80 / M tokens ≈ 60M tokens/month ≈ 2.0M ≈ 71M tokens/month
Frontier, from $60.35/mo ≈ $0.48 / M tokens ≈ 126M tokens/month ≈ 4.2M ≈ 148M tokens/month
Core, from $18.70/mo ≈ $0.17 / M tokens ≈ 111M tokens/month ≈ 3.7M ≈ 131M tokens/month

The savings factor

Per-token cost of the same tokens divided by the flat price (annual-billed anchor), per pool. Three usage profiles. The last column is a theoretical maximum: the arithmetic of the per-token line at that volume, not a level of service to expect — see "What unlimited and fair use mean here" below.

PoolLight use
60M tokens/mo
Daily driver
450M tokens/mo
Heavy agent loops
1.5B tokens/mo
Flagship ≈ 1.0×≈ 7.4×≈ 25× (theoretical max)
Frontier ≈ 0.5×≈ 3.6×≈ 12× (theoretical max)
Core ≈ 0.5×≈ 4.0×≈ 13× (theoretical max)

No SLA. Throughput, latency and the savings factor are not guaranteed: they depend on your workload mix and on the capacity available in the pool at the time. Plan with the daily-driver column; treat the theoretical maximum as an upper bound of the arithmetic, not as an expectation.

For scale: a single coding-agent task typically burns 300–500K tokens, because the agent re-sends its growing context on every tool call, and we measured a coding agent at roughly 2M tokens per hour. One to two million tokens a day is a handful of tasks, not a heavy day. Most people who run an agent daily are past break-even in the first week of the month.

What “unlimited” and fair use mean here

Section titled “What “unlimited” and fair use mean here”

Read the two tables with the right expectations, because a flat block is a different product from a meter:

  • Unlimited means tokens. Nothing is metered or billed per token during your reserved hours, and there are no overage charges. That is what makes the factor grow with your volume instead of the bill.
  • One request at a time per key. Throughput comes from concurrency, and concurrency comes from subscriptions: fold several into a combined key and parallel capacity stacks where their hours overlap.
  • Pools are shared, so fair use applies. When one key’s usage over a billing period runs far above what a single subscriber’s workload normally represents, the platform moderates that key’s speed — depending on how much capacity the pool has available at the time — so the pool keeps working for everyone. It slows; it does not cut off, and it never bills. The calibration is internal and not published; the rule itself is in our Terms, section 1.2.

Why we do not publish a sharper number. What a request costs to serve is not a token count: it depends on how much of the input is fresh versus served from cache, on request size, on output length, and on how loaded the pool is at that moment. The same million tokens can cost several times more or less depending on that mix. Moderating a key is also only one of several levers — capacity, pricing and moderation are balanced together, and that balance is the product. A single published threshold would be wrong for most workloads and right only for whoever tuned against it, so we publish the rule, not a figure.

So the columns read like this:

  • Light use is where per-token still competes. Below break-even, pay per token.
  • Daily driver is what an active developer or a busy agent actually reaches on one key: several times what the same month would cost per token. This is the column to plan with.
  • Heavy agent loops is the theoretical maximum: what the per-token line says at that volume, not a level of service you should expect. A key may run at that pace while the pool has headroom, and moderation applies as availability tightens, so treat anything above the daily-driver factor as upside rather than as the baseline — or buy it outright with more subscriptions, each with its own curve.
  • There is no SLA. We do not guarantee throughput, latency or a savings factor; the figures here are arithmetic on published prices and illustrative volumes, and the Terms govern.

That is the whole model in one sentence: flat rate wins by a factor that grows with your volume, up to what one key can push through its reserved block.

The curve is symmetric, so it also says when not to subscribe:

  • Sporadic use. A weekly batch job, an internal tool a few people touch once a day, a prototype you poke at on weekends. Below the break-even volume, per-token is cheaper and you should pay per token.
  • Usage concentrated outside your block. A block is a daily 8-hour window. If your traffic is spread evenly around the clock, you either need more blocks or you are paying for hours you do not use. Combined keys cover this: coverage windows add up across subscriptions.
  • Very high cache hit rates on a cheap model. Per-token vendors discount cached input heavily. A workload that is mostly repeated context on the cheapest model moves its break-even to the right. Run your own numbers with your cache rate before deciding.

If someone asks “is the subscription worth it for us?”, the answer is a volume, not an opinion:

  1. Estimate tokens per month per person or per agent. Count context re-sends; they dominate.
  2. Compare with the break-even column for the pool you would use.
  3. If you are past it, the flat block is cheaper by the factor in the second table, capped at what one key can push through. If you are not, pay per token.

Do these break-even volumes include cache discounts? No. They use the vendors’ list prices with no cache credit, which is the most favourable case for per-token pricing. With cache discounts, per-token gets cheaper and the break-even moves higher.

Is the savings factor guaranteed? No. It depends on your volume and on the throughput one key delivers. The daily-driver column is what a typical active developer sees; the heavy-loop column is the theoretical maximum, reachable while the pool has headroom and moderated under fair use as availability tightens. Subscriptions are unlimited in tokens, not in throughput.

What if I need more than one key can deliver? Take out additional subscriptions and fold them into a combined key. Where their hours overlap, parallel capacity stacks; where they do not, coverage extends. Each subscription is its own flat block with its own curve.

Is there an SLA? No. We do not guarantee throughput, latency or any savings factor. Fair-use moderation is not a service failure and does not give rise to refunds or credits; the Terms govern. The figures in this report are arithmetic on published prices and illustrative volumes.

Where are the current prices? Always on the pools page. This report recomputes from them on every build; the page is the source of truth.

Each pool is compared against the published per-token list price of the best comparable alternative for that class of model — the conservative case. Per-token prices are vendor list prices for the variants Artificial Analysis evaluates (peak list price where a vendor publishes peak/off-peak); open-weight models are often cheaper on aggregators, so the per-token side is conservative in the vendor’s favour. The 80/20 input/output mix is typical of agent workloads; a chat-heavy mix has more output and a higher blended rate. Usage profiles are illustrative volumes, not measurements of any subscriber.

  • 2026-09-17 — First edition: break-even volume and savings factor per pool, recomputed from live prices and the monthly Pareto edition.