Prompt caching cost: calculate reads, writes and real savings
Calculate cached input without double counting, include cache writes and storage where billed, and test whether normal reuse actually lowers task cost.
Published 2026-09-13 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

Prompt caching can lower the cost of repeated context, but a cache-hit percentage alone does not tell you whether a request became cheaper. Calculate uncached input, cache reads, any cache writes or storage, and output separately. Then compare the total cost of completing the same task. This guide shows the arithmetic and a small experiment you can repeat using your own usage records.
Prompt caching and response caching solve different problems
Prompt caching reuses the processing of eligible input while the model still generates a new answer. Response caching returns a previously stored answer. A support assistant can reuse a product manual as prompt context while answering a new question; returning yesterday's full answer may be wrong for today's order status. Keep the two measurements separate.
OpenAI's prompt caching guide explains prefix matching and model-specific cache behavior. Anthropic's documentation distinguishes cache creation from cache reads. Settings, minimum lengths and expiry differ. An OpenAI-compatible URL does not establish that every native cache option is available on a gateway route.
A worked request: avoid charging cached tokens twice
The following rates are illustrative assumptions, not a KeepRouter or provider quote. Assume a model reports 10,000 total input tokens, of which 8,000 are cached, plus 500 output tokens. Let ordinary input cost $2 per million tokens, cache reads $0.20, and output $8. This example has no separately billed cache write or storage.
| Component | Quantity | Calculation | Cost |
|---|---|---|---|
| Uncached input | 2,000 tokens | 2,000 × $2 / 1,000,000 | $0.0040 |
| Cache reads | 8,000 tokens | 8,000 × $0.20 / 1,000,000 | $0.0016 |
| Output | 500 tokens | 500 × $8 / 1,000,000 | $0.0040 |
| Total | One request | Sum of these three rows | $0.0096 |
Without a cache read, the same token quantities would cost $0.0240. The $0.0144 difference is 60% of that request's total, although cached input itself is priced 90% below ordinary input. Neither percentage is a measured saving for a real application. If your API reports uncached input separately, use that field directly instead of subtracting cached tokens again.
Include the first write and the unused cache
Suppose a reusable prefix would cost $0.02 at the ordinary input rate, writing it costs $0.025, and each read costs $0.002. These are another set of hypothetical prices. Two uses cost $0.027 with one write and one read, versus $0.04 without caching. One use costs more: $0.025 versus $0.02. Output, changing input and storage are excluded from this prefix-only example.
For N uses, compare write_cost + (N - 1) × read_cost + storage_cost with N × ordinary_prefix_cost. Use the selected model's actual pricing terms. Some APIs expose explicit caching, some implicit caching, and some both; Gemini's caching documentation describes its own mechanisms. Do not transfer a storage or write assumption across models without checking it.
Run a small cache experiment
Choose one model and endpoint from the live catalog. Save the displayed customer rates and date. Use a non-sensitive reference document and a short, fixed question, with a modest output limit. Confirm the document meets that model's cache requirements; padding a short prompt solely to reach a threshold can increase total cost.
Run the first request, repeat it, then change only the question after the shared context. Record input, cache-read, cache-write and output quantities when the endpoint returns them, plus elapsed time and the charge in your Usage view. Missing cache fields mean unknown cache usage, not a proven zero. A second identical request is not guaranteed to hit: eligibility, expiry, routing and prefix changes can all matter.
Finally repeat after the normal pause between real user requests. A hot loop can exaggerate reuse that will not survive your application's traffic pattern. Keep request identifiers and numeric usage; avoid putting document text or API keys into the experiment log.
Diagnose the result and choose the next change
| Observation | Likely investigation | Useful next step |
|---|---|---|
| Reads stay at zero | Eligibility, prefix changes, expiry or unsupported settings | Check the exact model's caching documentation |
| Reads rise but total cost does not fall | Writes, longer prompts, output or retries dominate | Compare all billable quantities per completed task |
| Estimated and recorded charges differ | Units, usage-field semantics or omitted components | Recalculate one isolated request before extrapolating |
| Short repeats save, normal sessions do not | Reuse interval is too long | Shorten context or change the caching strategy |
Use the API cost calculator for a scenario estimate and the cost-control guide for task-level budgeting. For separate Claude write/read accounting, use the Claude cache worksheet. Your next decision should follow the complete request cost and answer quality, rather than the cache-hit badge.
Frequently asked questions
Does an 80% cache-hit rate mean 80% lower cost?
No. Output, uncached input, cache writes, storage and retries can still be billable. Compare the total cost per completed task, using the selected model’s rates.
Should cached tokens be added to total input tokens?
Only if that API defines separate, non-overlapping quantities. When total input already includes cached tokens, subtract the cached portion before pricing ordinary input.
Sources reviewed
Article last reviewed 2026-09-29
- [1] OpenAI prompt caching
- [2] Anthropic prompt caching
- [3] Gemini context caching
- [4] KeepRouter customer model prices