# Prompt caching cost: calculate reads, writes and real savings

> Calculate cached input without double counting, include cache writes and storage where billed, and test whether normal reuse actually lowers task cost.

_Published 2026-09-13 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 4 minute read_

![Cache pricing audit diagram: a short green bar beside a tall accent bar, measured against the same baseline, with a ledger grid on the right](https://keeprouter.com/editorial/blog/llm-cache-pricing-audit.png)

_Compare the complete request cost with and without cache reuse. Illustration, not measured savings._

Prompt caching can lower the cost of repeated context, but a cache-hit percentage alone does not tell you whether a request became cheaper. Calculate uncached input, cache reads, any cache writes or storage, and output separately. Then compare the total cost of completing the same task. This guide shows the arithmetic and a small experiment you can repeat using your own usage records.

## Prompt caching and response caching solve different problems

Prompt caching reuses the processing of eligible input while the model still generates a new answer. Response caching returns a previously stored answer. A support assistant can reuse a product manual as prompt context while answering a new question; returning yesterday's full answer may be wrong for today's order status. Keep the two measurements separate.

[OpenAI's prompt caching guide](https://developers.openai.com/api/docs/guides/prompt-caching) explains prefix matching and model-specific cache behavior. [Anthropic's documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) distinguishes cache creation from cache reads. Settings, minimum lengths and expiry differ. An OpenAI-compatible URL does not establish that every native cache option is available on a gateway route.

## A worked request: avoid charging cached tokens twice

The following rates are **illustrative assumptions, not a KeepRouter or provider quote**. Assume a model reports 10,000 total input tokens, of which 8,000 are cached, plus 500 output tokens. Let ordinary input cost $2 per million tokens, cache reads $0.20, and output $8. This example has no separately billed cache write or storage.

| Component | Quantity | Calculation | Cost |
| --- | --- | --- | --- |
| Uncached input | 2,000 tokens | 2,000 × $2 / 1,000,000 | $0.0040 |
| Cache reads | 8,000 tokens | 8,000 × $0.20 / 1,000,000 | $0.0016 |
| Output | 500 tokens | 500 × $8 / 1,000,000 | $0.0040 |
| Total | One request | Sum of these three rows | $0.0096 |

Without a cache read, the same token quantities would cost $0.0240. The $0.0144 difference is 60% of that request's total, although cached input itself is priced 90% below ordinary input. Neither percentage is a measured saving for a real application. If your API reports uncached input separately, use that field directly instead of subtracting cached tokens again.

## Include the first write and the unused cache

Suppose a reusable prefix would cost $0.02 at the ordinary input rate, writing it costs $0.025, and each read costs $0.002. These are another set of hypothetical prices. Two uses cost $0.027 with one write and one read, versus $0.04 without caching. One use costs more: $0.025 versus $0.02. Output, changing input and storage are excluded from this prefix-only example.

For N uses, compare `write_cost + (N - 1) × read_cost + storage_cost` with `N × ordinary_prefix_cost`. Use the selected model's actual pricing terms. Some APIs expose explicit caching, some implicit caching, and some both; [Gemini's caching documentation](https://ai.google.dev/gemini-api/docs/caching) describes its own mechanisms. Do not transfer a storage or write assumption across models without checking it.

## Run a small cache experiment

Choose one model and endpoint from the [live catalog](/models). Save the displayed customer rates and date. Use a non-sensitive reference document and a short, fixed question, with a modest output limit. Confirm the document meets that model's cache requirements; padding a short prompt solely to reach a threshold can increase total cost.

Run the first request, repeat it, then change only the question after the shared context. Record input, cache-read, cache-write and output quantities when the endpoint returns them, plus elapsed time and the charge in your Usage view. Missing cache fields mean unknown cache usage, not a proven zero. A second identical request is not guaranteed to hit: eligibility, expiry, routing and prefix changes can all matter.

Finally repeat after the normal pause between real user requests. A hot loop can exaggerate reuse that will not survive your application's traffic pattern. Keep request identifiers and numeric usage; avoid putting document text or API keys into the experiment log.

## Diagnose the result and choose the next change

| Observation | Likely investigation | Useful next step |
| --- | --- | --- |
| Reads stay at zero | Eligibility, prefix changes, expiry or unsupported settings | Check the exact model's caching documentation |
| Reads rise but total cost does not fall | Writes, longer prompts, output or retries dominate | Compare all billable quantities per completed task |
| Estimated and recorded charges differ | Units, usage-field semantics or omitted components | Recalculate one isolated request before extrapolating |
| Short repeats save, normal sessions do not | Reuse interval is too long | Shorten context or change the caching strategy |

Use the [API cost calculator](/tools/api-cost-calculator) for a scenario estimate and the [cost-control guide](/blog/control-multi-model-api-costs) for task-level budgeting. For separate Claude write/read accounting, use the [Claude cache worksheet](/blog/claude-api-cost-prompt-caching). Your next decision should follow the complete request cost and answer quality, rather than the cache-hit badge.

## Frequently asked questions

### Does an 80% cache-hit rate mean 80% lower cost?

No. Output, uncached input, cache writes, storage and retries can still be billable. Compare the total cost per completed task, using the selected model’s rates.

### Should cached tokens be added to total input tokens?

Only if that API defines separate, non-overlapping quantities. When total input already includes cached tokens, subtract the cached portion before pricing ordinary input.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [OpenAI prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching)
2. [Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
3. [Gemini context caching](https://ai.google.dev/gemini-api/docs/caching)
4. [KeepRouter customer model prices](https://keeprouter.com/models)

## Related guides

- [How to reduce LLM API costs: a task-level cost worksheet](https://keeprouter.com/blog/control-multi-model-api-costs.md)
- [LLM routing vs load balancing: four policies teams often confuse](https://keeprouter.com/blog/llm-routing-vs-load-balancing.md)
- [kimi k2.7 code highspeed](https://keeprouter.com/models/kimi-k2.7-code-highspeed.md)
- [quickstart](https://keeprouter.com/docs/quickstart.md)

## Put your workload into the cost estimate

Choose a model and enter expected usage. Compare the estimate with a small real request before scaling.

[Estimate API costs](https://keeprouter.com/tools/api-cost-calculator)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
