# Claude API cost: plan prompt caching around writes, reads and reuse

> Claude caching can change the cost of a reused prompt prefix, but cache creation and cache reads are different operations. Budget the first request and later requests separately, keep output costs, and use the billing rules of the endpoint you call.

_Published 2026-09-23 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Multi-model API cost ledger showing tokens, cache, retries, fallbacks, and ownership](https://keeprouter.com/editorial/blog/control-multi-model-api-costs.png)

_Cache writes and later reads belong in separate rows of the cost worksheet._

Claude API cost is easier to understand when a conversation is treated as a sequence of requests. A large instruction or reference document may be written to a prompt cache once and read on later requests. New questions and generated answers still add tokens. A cache discount applied to the entire conversation therefore produces a misleading budget.

Use this guide for a document review assistant that asks several questions about the same material. It focuses on workload arithmetic. Anthropic's [prompt caching documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) defines its native cache behavior; the [KeepRouter Claude Sonnet 4.6 page](/models/claude-sonnet-4-6) supplies the current customer rate and endpoint for that catalog route. This article does not claim that every gateway preserves every native caching option.

## Name the four costs on the worksheet

| Bucket | What causes it | What to record |
|---|---|---|
| Ordinary input | New material outside the reused prefix | Input tokens and applicable rate |
| Cache creation | A prefix is written for later reuse | Creation tokens and write duration |
| Cache read | A reusable prefix is found | Read tokens and read rate |
| Output | A new answer is generated | Output tokens and output rate |

Anthropic documents cache creation and cache reads separately. Cache lifetime, supported model, minimum useful prefix length, and the exact content preceding a breakpoint affect whether reuse is possible. Keep stable instructions and reference material ahead of the changing question. Do not put a random request identifier at the beginning of a prefix you intend to reuse.

For a gateway request, verify the public contract and charge rather than importing Anthropic's retail multipliers automatically. KeepRouter publishes input, cached-input and output rates. Its current billing normalization folds native cache-creation tokens into billable input and records cache-read tokens separately. Reconcile the route you use with the console; an upstream cache-write price does not by itself establish a separate KeepRouter customer surcharge.

## Calculate when a cache write pays for itself

Let P be the reused prefix size, N the number of requests, B the ordinary input rate, W the write rate, and R the read rate, all priced per million tokens. Without reuse, the prefix costs `N * P * B / 1_000_000`. With one write and N minus one reads, it costs `P * (W + (N - 1) * R) / 1_000_000`.

Suppose, only for arithmetic, B is $2, W is $2.50, R is $0.20, P is 10,000 tokens and N is 3. The uncached prefix costs $0.06. One write and two reads cost $0.029. Add the changing questions and all generated answers to both scenarios before comparing totals. These numbers are hypothetical; they are not a price quote or a measured saving.

The break-even condition is `W + (N - 1) * R < N * B`. If B is greater than R, this can be rearranged to `N > (W - R) / (B - R)`. Actual expiration or prefix changes can create extra writes, so a session with three turns may not behave like one write plus two reads. Record what happened before using the formula in a budget.

## Inspect usage without exposing document content

For a native Messages response, preserve the distinction between ordinary input, created cache, and cache reads. The following Python fragment reads a saved response object; it makes no network call and prints no prompt or API key.

```python
def claude_usage(response):
    usage = response["usage"]
    required = ("input_tokens", "output_tokens")
    if any(name not in usage for name in required):
        raise ValueError("Missing usage; inspect request record")
    return {
        "ordinary_input": usage["input_tokens"],
        "cache_created": usage.get("cache_creation_input_tokens", 0),
        "cache_read": usage.get("cache_read_input_tokens", 0),
        "output": usage["output_tokens"],
    }
```

Do not use this parser on an OpenAI Chat Completions response without adapting its semantics. In that format cached input is generally included in total prompt tokens; the native Messages fields keep categories separate. If you normalize both formats into a spreadsheet, name the final columns consistently and keep a sample raw object for each route.

## Run a small reuse experiment

Choose a synthetic document long enough for the selected model's current cache requirements. Send the same stable prefix with three distinct questions using the documented native cache settings for the tested route. Record creation, read and output counters on each request, plus when the request started. Repeat once after the intended idle interval. This distinguishes a useful conversational cache from a cache that expires before the user returns.

Then test your actual prompt builder. Changing the tool definitions, moving a system instruction, or inserting metadata may change the reusable prefix. A useful engineering outcome is a fixture showing the exact serialized prefix that remains constant. Avoid reporting a percentage saving until the charges and completed task quality have been measured together.

## Choose cache lifetime from the reuse pattern

Compare repeated requests at the intervals your application actually sees, not only back-to-back calls. Record creation and read quantities separately, and include unused writes in the total. [Anthropic’s cache documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) defines supported lifetimes and pricing behavior. If long pauses prevent reuse, a shorter prompt may be a better cost experiment than a larger cache. Use the [cache worksheet](/blog/llm-cache-pricing-audit) to check the break-even arithmetic.

## Choose a budget for the complete workflow

Use the [API cost calculator](/tools/api-cost-calculator) for the published input, cached-input and output buckets. For direct Anthropic calls with separate write prices, keep the write calculation in your worksheet rather than squeezing it into a read-only cache input. Include retries, answers discarded by validation, and the first cold request after a long idle period.

For interactive document review, compare three schedules: every question arrives immediately, users pause between questions, and each document is used once. A cache strategy that works for the first schedule may do little for the last. Reduce unnecessary document repetition and cap answer length where the task allows it, then check whether the assistant still answers correctly.

Keep the [cache pricing audit](/blog/llm-cache-pricing-audit) beside the cost worksheet. If your workload runs inside Claude Code, the [Claude Code cost comparison guide](/blog/cut-claude-code-costs) adds task completion and tool behavior to the evaluation. A lower prefix charge alone does not establish a lower cost per finished task.

## Frequently asked questions

### Are cache writes and reads billed identically?

Not necessarily. Native Anthropic pricing distinguishes them. A gateway may normalize usage differently, so use its published customer rates and recorded charge.

### Does prompt caching reuse the previous answer?

No. It reuses prompt processing; generated output remains part of each request.

### Should cached input include cache-creation tokens?

Keep writes and reads separate in the worksheet. Use the response format and billing service to decide how each maps to a charge.

### Can a short test prove production savings?

It can establish observed behavior for that sequence. Production cost also depends on reuse frequency, idle gaps, retries and completed-task quality.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)
2. [Anthropic Messages reference](https://platform.claude.com/docs/en/api/messages/create)
3. [KeepRouter cache pricing audit](https://keeprouter.com/blog/llm-cache-pricing-audit)

## Related guides

- [claude sonnet 4 6](https://keeprouter.com/models/claude-sonnet-4-6.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)
- [Prompt caching cost: calculate reads, writes and real savings](https://keeprouter.com/blog/llm-cache-pricing-audit.md)
- [Compare Claude Code model costs without a static price snapshot](https://keeprouter.com/blog/cut-claude-code-costs.md)
- [DeepSeek API pricing: estimate cache hits without counting input twice](https://keeprouter.com/blog/deepseek-api-pricing-cache-estimation.md)
- [KeepRouter API key setup: from the free model to a controlled paid request](https://keeprouter.com/blog/keeprouter-api-key-free-to-paid.md)

## Put your workload into the cost estimate

Choose a model and enter expected usage. Compare the estimate with a small real request before scaling.

[Estimate API costs](https://keeprouter.com/tools/api-cost-calculator)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
