2026-09-30
Your model price table is lying to you
Cache multipliers vary 20x across models, and most cost tools treat them as a constant. The number stays plausible, which is why nobody catches it.
Every LLM cost dashboard I have looked at gets one thing structurally wrong, and it is not a rounding error. It treats prompt caching as a constant multiplier: cache reads cost 10% of input, everywhere, for every model.
That was roughly true in 2024. In 2026 it is wrong by up to 20x — and the error is concentrated in exactly the line item that caching was supposed to shrink, so it stays plausible.
What the providers actually charge
Real cache-read multipliers, read from the providers' own pricing pages:
| Model | Input /1M | Cache read /1M | Multiplier |
|---|---|---|---|
| Claude Fable 5.1 | $10.00 | $0.25 | 0.025x |
| Claude Opus 5.5 | $4.00 | $0.20 | 0.05x |
| gpt-6.1-sol | $2.00 | $0.10 | 0.05x |
| Claude Sonnet 5.5 | $2.00 | $0.20 | 0.1x |
| gpt-5.5 | $5.00 | $0.50 | 0.1x |
| gpt-4o | $2.50 | $1.25 | 0.5x |
Six models, five different multipliers, spanning a 20x range.
Why the mistake survives
Look at the last row. gpt-4o is the one model where 0.1 is not quite the right answer either — its real cache multiplier is 0.5.
The constant was almost certainly calibrated years ago against an early model and then never revisited. Nobody has a reason to question it, because the output never looks broken: the total spend figure is dominated by fresh input and output tokens, so a wrong cache rate shifts the headline number by a few percent.
The error hides in the line item that is small enough not to check.
Where it actually hurts
Take a realistic agent workload: 30M input tokens a month, 60% of them cache reads (18M cached, 12M fresh). Here is the cache-read line item alone, computed with a hardcoded 0.1 versus the real rate:
| Model | Assumed (0.1x) | Real | Error |
|---|---|---|---|
| gpt-4o | $4.50 | $22.50 | understated 5x |
| Claude Opus 5.5 | $7.20 | $3.60 | overstated 2x |
| gpt-6.1-sol | $3.60 | $1.80 | overstated 2x |
| Claude Fable 5.1 | $18.00 | $4.50 | overstated 4x |
| Claude Sonnet 5.5 | $3.60 | $3.60 | correct |
To be precise about the blast radius: on total spend across this workload, the error lands in the 5-15% range, because fresh input and output tokens dominate. That is small enough to explain why nobody notices, and large enough to matter when you are trying to decide whether prompt caching paid for itself.
The problem is not that the total is wildly wrong. It is that the number you would use to evaluate caching is wrong by 2x to 5x, in both directions depending on which model you are on.
Cache writes are a billed category, not a rounding detail
The second structural miss: writing to the cache is charged, at a premium.
- OpenAI bills cache writes on the gpt-6 family —
gpt-6-astracharges $12.50/1M to write against $10.00/1M for plain input. - Anthropic charges 1.25x input for a 5-minute cache write, and 2x for a one-hour write.
If your schema has three token columns (input, output, total), you cannot represent this. You need four: input, cached input, cache write, output. That schema change is the difference between a plausible report and a correct one.
A third miss: context length changes the rate
Current OpenAI flagships bill prompts over 272K tokens at roughly 2x the short-context rate. gpt-6-astra goes from $10 to $20 input, and $50 to $75 output.
A single blended per-token rate cannot express that either. If you route long-context work through a cost tool that does not model tiers, you will understate it by up to half.
The failure mode we hit ourselves
When we first wrote the price table for llm-guard, we filled it with the models we knew: gpt-4o, claude-3-5-sonnet, o3. Then we opened the providers' pricing pages for real.
| What we had | What is current | |
|---|---|---|
| OpenAI | gpt-4o, gpt-4.1, o1/o3 | gpt-6-astra, gpt-6.1-sol, gpt-6-luna |
| Anthropic | claude-3-5-sonnet, claude-3-opus | Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 4.5 |
Our table was about a year out of date. Every current model resolved as unpriced, and the tool was reporting $0.00 for the models people actually use.
Two structural decisions came out of that:
- Unpriced is a first-class state. A model with no price on file gets
cost_usd = NULLand shows up in the report as "N models have no price entry." We refuse to fall back to zero. If we had, that stale table would have quietly reported a real month as free — which is a far worse failure than the cache-multiplier bug above. - The table carries a verification date and a documented refresh procedure, so its staleness is auditable rather than invisible.
How to check your own numbers
You do not need our tool for this. Three questions:
# 1. Does your tool report a cost at all for a model released this year?
# 2. Ask it what a cache read costs on your most expensive model.
# 3. Ask it what a cache WRITE costs.
If the answer to (2) is the same number for every model, or the answer to (3) is "we don't track that," you have found the bug.
Run the numbers on your own bill
We turned this into a calculator. Pick your model, enter your monthly token volume and cache hit rate, and it shows the real cache-read price for your model, how far off the flat 10% assumption is, and what raising your hit rate is worth.
Every figure verified against the providers' pricing pages on 2026-09-30. The full table, the sources and the quarterly refresh procedure are in llm-guard's pricing documentation.