Prompt caching cost: calculate reads, writes and real savings

Calculate cached input without double counting, include cache writes and storage where billed, and test whether normal reuse actually lowers task cost.

Published 2026-09-13 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

Cache pricing audit diagram: a short green bar beside a tall accent bar, measured against the same baseline, with a ledger grid on the right
Compare the complete request cost with and without cache reuse. Illustration, not measured savings.

Prompt caching can lower the cost of repeated context, but a cache-hit percentage alone does not tell you whether a request became cheaper. Calculate uncached input, cache reads, any cache writes or storage, and output separately. Then compare the total cost of completing the same task. This guide shows the arithmetic and a small experiment you can repeat using your own usage records.

Prompt caching and response caching solve different problems

Prompt caching reuses the processing of eligible input while the model still generates a new answer. Response caching returns a previously stored answer. A support assistant can reuse a product manual as prompt context while answering a new question; returning yesterday's full answer may be wrong for today's order status. Keep the two measurements separate.

OpenAI's prompt caching guide explains prefix matching and model-specific cache behavior. Anthropic's documentation distinguishes cache creation from cache reads. Settings, minimum lengths and expiry differ. An OpenAI-compatible URL does not establish that every native cache option is available on a gateway route.

A worked request: avoid charging cached tokens twice

The following rates are illustrative assumptions, not a KeepRouter or provider quote. Assume a model reports 10,000 total input tokens, of which 8,000 are cached, plus 500 output tokens. Let ordinary input cost $2 per million tokens, cache reads $0.20, and output $8. This example has no separately billed cache write or storage.

ComponentQuantityCalculationCost
Uncached input2,000 tokens2,000 × $2 / 1,000,000$0.0040
Cache reads8,000 tokens8,000 × $0.20 / 1,000,000$0.0016
Output500 tokens500 × $8 / 1,000,000$0.0040
TotalOne requestSum of these three rows$0.0096

Without a cache read, the same token quantities would cost $0.0240. The $0.0144 difference is 60% of that request's total, although cached input itself is priced 90% below ordinary input. Neither percentage is a measured saving for a real application. If your API reports uncached input separately, use that field directly instead of subtracting cached tokens again.

Include the first write and the unused cache

Suppose a reusable prefix would cost $0.02 at the ordinary input rate, writing it costs $0.025, and each read costs $0.002. These are another set of hypothetical prices. Two uses cost $0.027 with one write and one read, versus $0.04 without caching. One use costs more: $0.025 versus $0.02. Output, changing input and storage are excluded from this prefix-only example.

For N uses, compare write_cost + (N - 1) × read_cost + storage_cost with N × ordinary_prefix_cost. Use the selected model's actual pricing terms. Some APIs expose explicit caching, some implicit caching, and some both; Gemini's caching documentation describes its own mechanisms. Do not transfer a storage or write assumption across models without checking it.

Run a small cache experiment

Choose one model and endpoint from the live catalog. Save the displayed customer rates and date. Use a non-sensitive reference document and a short, fixed question, with a modest output limit. Confirm the document meets that model's cache requirements; padding a short prompt solely to reach a threshold can increase total cost.

Run the first request, repeat it, then change only the question after the shared context. Record input, cache-read, cache-write and output quantities when the endpoint returns them, plus elapsed time and the charge in your Usage view. Missing cache fields mean unknown cache usage, not a proven zero. A second identical request is not guaranteed to hit: eligibility, expiry, routing and prefix changes can all matter.

Finally repeat after the normal pause between real user requests. A hot loop can exaggerate reuse that will not survive your application's traffic pattern. Keep request identifiers and numeric usage; avoid putting document text or API keys into the experiment log.

Diagnose the result and choose the next change

ObservationLikely investigationUseful next step
Reads stay at zeroEligibility, prefix changes, expiry or unsupported settingsCheck the exact model's caching documentation
Reads rise but total cost does not fallWrites, longer prompts, output or retries dominateCompare all billable quantities per completed task
Estimated and recorded charges differUnits, usage-field semantics or omitted componentsRecalculate one isolated request before extrapolating
Short repeats save, normal sessions do notReuse interval is too longShorten context or change the caching strategy

Use the API cost calculator for a scenario estimate and the cost-control guide for task-level budgeting. For separate Claude write/read accounting, use the Claude cache worksheet. Your next decision should follow the complete request cost and answer quality, rather than the cache-hit badge.

Frequently asked questions

Does an 80% cache-hit rate mean 80% lower cost?

No. Output, uncached input, cache writes, storage and retries can still be billable. Compare the total cost per completed task, using the selected model’s rates.

Should cached tokens be added to total input tokens?

Only if that API defines separate, non-overlapping quantities. When total input already includes cached tokens, subtract the cached portion before pricing ordinary input.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] OpenAI prompt caching
  2. [2] Anthropic prompt caching
  3. [3] Gemini context caching
  4. [4] KeepRouter customer model prices

Related guides

← All posts · Models & pricing · Get an API key