# Reasoning Token Costs: Read Usage Without Double Counting

> Understand reasoning-token billing across API contracts. Reconcile visible output, usage details and failed attempts with a reproducible cost example.

_Published 2026-09-29 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Multi-model API cost ledger showing tokens, cache, retries, fallbacks, and ownership](https://keeprouter.com/editorial/blog/control-multi-model-api-costs.png)

_Compare the full cost of accepted tasks with explicit workload assumptions. Illustration, not a vendor quote._

A short answer can still have substantial generation cost when a model uses reasoning tokens. Read the usage contract for the exact API and model: some fields are components of a larger total, while other APIs report categories that must be added. Counting every numeric field as a separate charge can overstate the bill.

This guide explains the accounting decision. It does not promise identical reasoning controls, usage fields or billing behavior across OpenAI, Gemini and a gateway that exposes compatible requests.

## Start with the native accounting rules

[OpenAI's reasoning documentation](https://developers.openai.com/api/docs/guides/reasoning) describes reasoning tokens as billed output and part of context consumption. A reasoning breakdown is therefore not automatically an extra category to add on top of the output total that already contains it.

[Google's current thinking guide](https://ai.google.dev/gemini-api/docs/thinking) describes Interactions API fields `usage.total_output_tokens` and `usage.total_thought_tokens`, with response pricing based on output plus thinking. Those categories are additive under that documented contract; they are not the OpenAI-style nested breakdown in the example below. Do not transplant a formula written for one API into another merely because both include a field named “output.”

Before building a dashboard, write a small mapping for the response format you actually receive:

| Question | Why it matters |
| --- | --- |
| Is reasoning included in an output total? | Determines whether adding it again duplicates charges |
| Is a field a detail or an independent category? | Prevents summing totals and their components |
| Are cached tokens already included in input? | Changes how regular and cached rates apply |
| Which model and route produced this record? | Selects the correct rate and accounting contract |
| Did the request finish or fail after generation began? | A failed user task may still have measured usage |

Keep the original usage record available for reconciliation, with sensitive response content excluded. Normalize only after deciding what each field means.

## Reproduce a double-counting error

Consider a hypothetical OpenAI-style record with 1,000 total output tokens, of which 800 are reasoning tokens and 200 are visible answer tokens. At an illustrative $10 per million output tokens, output costs $0.01. Charging 1,800 tokens would incorrectly add the reasoning component twice.

```python
from decimal import Decimal

output_total = 1_000
reasoning_detail = 800
price_per_million = Decimal("10")
correct = Decimal(output_total) * price_per_million / 1_000_000
wrong = Decimal(output_total + reasoning_detail) * price_per_million / 1_000_000
assert correct == Decimal("0.01")
assert wrong == Decimal("0.018")
```

This example covers only output charges and uses a fictional price. It is not a universal formula for Gemini or another API. Input, caching, tools and other billed operations need their own contract-aware treatment. Do not infer a complete invoice from this single worksheet.

The opposite error is to count only visible text. In the same example that would count 200 tokens and miss most generated usage. Neither the displayed answer length nor a local tokenizer applied to that answer reconstructs hidden reasoning usage reliably.

## Set controls for the task, not for appearances

Where the chosen model and route expose reasoning controls, test supported effort or thinking settings on the same workload. Parameter names and accepted values vary. An OpenAI-compatible chat route does not itself prove that every native reasoning parameter is preserved or supported.

Choose a task with a checkable result. For example, ask the model to identify a contradiction between two fictional shipping rules: standard delivery excludes islands, while a later exception explicitly includes Island A. A passing answer must apply the exception and cite both rule IDs. Also include a simple lookup that needs no multi-step reasoning.

Compare accepted answers, total charged usage and elapsed time at each supported setting. More reasoning can be worthwhile for the contradiction task and wasteful for the lookup. A single global setting hides that difference. Keep the task mix constant during the comparison.

## Avoid an output cap that creates paid failures

Output limits and reasoning budgets interact differently across models. A restrictive generation cap may be consumed before the application gets a useful final answer. An empty or incomplete response is not evidence that no computation happened.

Record terminal status, usage and the application's acceptance decision separately. If a task fails, inspect whether the cause was insufficient generation room, unsupported parameters, missing context or a transport problem before increasing the budget. The [streaming guide](/blog/llm-streaming-interrupted) explains why some visible text is also not proof of completion.

A retry is a new attempt in the cost ledger. If the first attempt consumed $0.01 and the successful retry consumed $0.02, that task cost $0.03 before any other charges. Reporting only the successful attempt makes reasoning settings look cheaper than they are.

## Use the data to select a route

From the [model catalog](/models), choose candidates whose documented capabilities match the request type. Use the [API cost calculator](/tools/api-cost-calculator) for a planning estimate with supported price categories, then reconcile a bounded test against actual usage and charges. A calculator cannot recover a missing usage record or determine how an undocumented field is billed.

Compare cost per accepted task and latency at equivalent quality. Keep API name, model revision, route, reasoning setting and test date with the result. When any of those changes, rerun the fixture before treating the previous estimate as current. This gives you a defensible spending decision without turning a short answer into a misleading cost proxy.

## Frequently asked questions

### Are reasoning tokens billed even when they are not visible?

For models and APIs that bill reasoning, yes. Check the exact usage and pricing contract; visible answer length is not a complete measure of generation.

### Should I add reasoning tokens to output tokens?

Only if the API contract defines them as separate billable categories. If the output total already includes reasoning, adding the reasoning detail again double counts it.

### Does less reasoning always reduce task cost?

No. If acceptance falls and retries increase, the completed task can cost more. Compare all attempts on the same quality-controlled workload.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [OpenAI reasoning accounting](https://developers.openai.com/api/docs/guides/reasoning)
2. [Gemini thinking and pricing](https://ai.google.dev/gemini-api/docs/thinking)

## Related guides

- [RAG Cost per Query: A Worksheet Beyond Token Prices](https://keeprouter.com/blog/rag-api-cost-per-query.md)
- [LLM Streaming Stops Early: Diagnose SSE and Completion](https://keeprouter.com/blog/llm-streaming-interrupted.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)

## Put your workload into the cost estimate

Choose a model and enter expected usage. Compare the estimate with a small real request before scaling.

[Estimate API costs](https://keeprouter.com/tools/api-cost-calculator)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
