# DeepSeek API pricing: estimate cache hits without counting input twice

> A DeepSeek cost estimate needs three token buckets and the rates of the service you actually pay. A repeated prompt does not guarantee a cache hit. Read usage, separate cached input from total input, then reconcile the estimate with the billed request.

_Published 2026-09-23 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Cache pricing audit diagram: a short green bar beside a tall accent bar, measured against the same baseline, with a ledger grid on the right](https://keeprouter.com/editorial/blog/llm-cache-pricing-audit.png)

_Split uncached input, cached input and output before applying the current rates._

A useful DeepSeek API budget starts with uncached input, cached input, and generated output. Apply the rate for each bucket once. The current [KeepRouter model catalog](/models) supplies customer rates for its routes; DeepSeek's own pricing applies to calls billed by DeepSeek. A shared model name does not make those two bills interchangeable.

This guide builds an estimate for a document question-answering workload. The examples are arithmetic and implementation instructions, not production measurements. Start with the current [DeepSeek V4 Pro catalog entry](/models/deepseek-v4-pro), confirm its enabled endpoint, and select that KeepRouter model and enter your usage assumptions in the [API cost calculator](/tools/api-cost-calculator). For direct-provider prices, use the formula below in a separate worksheet. A catalog ID identifies a KeepRouter route; it is not evidence that every similarly named upstream deployment has the same capabilities.

## Split the usage before multiplying

For a Chat Completions usage object, total prompt tokens include cached prompt tokens. DeepSeek documents both `prompt_cache_hit_tokens` and `prompt_cache_miss_tokens`; compatible responses can also expose `prompt_tokens_details.cached_tokens`. Treat these as alternative representations of the same cache-hit count, not independent discounts.

| Variable | Meaning | How it enters the estimate |
|---|---|---|
| I | Total input tokens | Split into cached and uncached parts |
| H | Input tokens reported as cache hits | Multiply by the cached-input rate |
| O | Output tokens | Multiply by the output rate |
| Pi, Ph, Po | USD per million tokens | Use one service's price sheet |

The estimate is `((I - H) * Pi + H * Ph + O * Po) / 1_000_000`. Check that `0 <= H <= I`. Do not subtract cached input twice if your exported usage already calls its input column "uncached input". Keep raw response usage beside the normalized values so a future billing investigation can recover the original meaning.

## Work through a hypothetical document request

Suppose a request reports 20,000 input tokens, 16,000 cache-hit tokens, and 1,000 output tokens. For illustration only, choose rates of $1 per million uncached input, $0.10 cached input, and $4 output. These are invented round numbers for the calculation, not a DeepSeek or KeepRouter quote.

The uncached portion costs $0.004, the cached portion $0.0016, and output $0.004. The estimated total is $0.0096. With no cache hit, the same token counts would cost $0.024. This comparison holds output length and prices fixed. It does not predict the cache rate of your next request or the number of tokens another model will generate.

For a monthly budget, add requests individually or group requests with similar measured sizes. Do not multiply this favorable example by every monthly request. A document uploaded once and never reused has a different cost profile from a document queried repeatedly. Keep cold requests, repeated requests, retries, and long answers in separate rows.

## Parse a response with explicit field semantics

This parser accepts the common DeepSeek/OpenAI Chat Completions fields. It is a local accounting example, not a replacement for the service's billing ledger.

```python
def token_buckets(usage):
    total = usage.get("prompt_tokens")
    details = usage.get("prompt_tokens_details") or {}
    hit = details.get("cached_tokens")
    if hit is None:
        hit = usage.get("prompt_cache_hit_tokens", 0)
    if total is None:
        miss = usage.get("prompt_cache_miss_tokens")
        if miss is None:
            raise ValueError("No input breakdown; inspect raw usage")
        total = hit + miss
    output = usage.get("completion_tokens")
    values = (total, hit, output)
    if any(type(v) is not int or v < 0 for v in values):
        raise ValueError("Incomplete or invalid token counts")
    if hit > total:
        raise ValueError("Cache hits exceed total input")
    return {"uncached": total - hit, "cached": hit, "output": output}
```

A missing output count should remain an unresolved record, not silently become free output. Streaming clients must retain the final usage event when the route provides one. A client cancellation can interrupt collection; use the request's usage record to investigate rather than declaring the cancelled request unbilled.

## Design a cache experiment that answers one question

DeepSeek describes prefix caching as best effort. Its current rules distinguish persisted prefix units from arbitrary overlapping text. Therefore, "I sent nearly the same prompt twice" is not enough to prove a cache defect.

Choose one synthetic document and a fixed system instruction. Record the request order and completion times. Send a first question, then a follow-up conversation preserving the preceding messages, then a separate cold request with a different document. Record the cache counters for each. Change one factor per run: document prefix, message ordering, or spacing. Do not simultaneously change model, endpoint, and prompt layout.

Your result should say which specific sequence produced reported hits, with the model ID and date. A shorter response time is useful operational evidence but cannot establish a cache hit by itself. A cache hit also does not replay the previous answer: the model still generates output, which can differ.

## Use a fixture that catches double counting

A response fixture with 20,000 total input tokens and 16,000 cache-hit tokens should produce 4,000 ordinary-input tokens. Add another fixture where cache hits exceed total input and require an error instead of a negative charge. [DeepSeek’s cache documentation](https://api-docs.deepseek.com/guides/kv_cache/) explains the provider mechanism; interpret the fields actually returned by your chosen endpoint. This small arithmetic check is useful before connecting the estimate to a monthly budget.

## Turn the experiment into a budget

Use measured input and output distributions for the target feature. Keep a conservative zero-hit scenario, an observed-hit scenario, and a longer-answer scenario. If their cost difference is large, put a request or task cap in the application before broad rollout. Reconcile the first paid sample against the console's recorded charge and preserve any discrepancy for investigation.

For the broader accounting distinction, read the [cache pricing audit](/blog/llm-cache-pricing-audit). For code configuration, use the [OpenAI SDK guide](/use-cases/openai-sdk). Return to the calculator when prices or traffic change; do not preserve this example's invented rates as a production default.

## Frequently asked questions

### Does repeating a prompt guarantee a DeepSeek cache hit?

No. DeepSeek describes best-effort prefix caching. Inspect the reported cache counters for the exact request sequence.

### Should cached tokens be added to prompt_tokens?

Not when prompt_tokens already includes them. Split the total into uncached and cached input before pricing.

### Are the worked-example rates current prices?

No. They are hypothetical round numbers. Use the live price for the service billing your request.

### Can I estimate the whole month from one warm request?

That is unreliable. Separate cold requests, repeated contexts, retries and output lengths before projecting volume.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [DeepSeek context caching](https://api-docs.deepseek.com/guides/kv_cache/)
2. [DeepSeek Chat Completions usage fields](https://api-docs.deepseek.com/api/create-chat-completion/)
3. [KeepRouter current model prices](https://keeprouter.com/models)

## Related guides

- [deepseek v4 pro](https://keeprouter.com/models/deepseek-v4-pro.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)
- [Prompt caching cost: calculate reads, writes and real savings](https://keeprouter.com/blog/llm-cache-pricing-audit.md)
- [openai sdk](https://keeprouter.com/use-cases/openai-sdk.md)
- [Claude API cost: plan prompt caching around writes, reads and reuse](https://keeprouter.com/blog/claude-api-cost-prompt-caching.md)
- [GPT API pricing and SDK migration: keep the request contract visible](https://keeprouter.com/blog/gpt-api-pricing-sdk-migration.md)

## Put your workload into the cost estimate

Choose a model and enter expected usage. Compare the estimate with a small real request before scaling.

[Estimate API costs](https://keeprouter.com/tools/api-cost-calculator)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
