# Text-to-speech billing: input tokens plus audio seconds

> A speech request can have both input-token and audio-output charges. Separate the two rates, estimate duration, and reconcile the complete request rather than treating the audio rate as the whole bill.

_Published 2026-09-12 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Duration ruler with a waveform bar, a play node, and a measured span](https://keeprouter.com/editorial/blog/text-to-speech-per-second-billing.png)

_Measure returned audio duration and count input tokens separately; the selected model's published rates determine both charges._

**Short answer:** a KeepRouter speech request has two priced parts: text input and generated audio. The current `gemini-3.1-flash-tts-preview` route charges input tokens plus seconds of audio output. Looking only at the per-second rate leaves the input charge out of the estimate.

The [model page](/models/gemini-3.1-flash-tts-preview) and [API cost calculator](/tools/api-cost-calculator?model=gemini-3.1-flash-tts-preview) show the current customer prices. The worked numbers below were checked on September 23, 2026. They describe KeepRouter billing; a provider's direct invoice is a separate contract.

## Keep the two quantities separate

The request sends text and the response body is an audio file. For this route, KeepRouter estimates input tokens from the supplied text and measures or estimates the returned audio duration. A binary response does not carry a normal Chat Completions usage object, so keep the console request record when reconciling a charge.

Two texts of equal length can produce different durations because of pacing, punctuation, and voice. Input length alone therefore cannot predict the whole bill. Conversely, audio duration alone does not include the text-input charge. The estimate is:

```
cost = input tokens x input price per token
     + generated audio seconds x output price per second
```

Use the quantities recorded for the actual request when checking a bill. Before calling, treat token count and duration as assumptions rather than guaranteed output. An input token estimate is not the same measurement as the number of characters or words in the source text.

## Where the audio rate comes from

Google's Gemini pricing page states 25 audio tokens per second for this TTS model. Its standard audio-output price of $20 per million audio tokens converts to `$20 x 25 / 1,000,000 = $0.0005` per second. That conversion applies to the audio-output part only; it does not remove the separate text-input price.

KeepRouter's published customer rates at this review are **$1 per million input tokens** and **$0.0005 per second of generated audio**. Always recheck the catalog before budgeting. Provider rates help explain the conversion, but the KeepRouter model page is the price source for a KeepRouter request.

The earlier shorthand of one cent for twenty seconds describes **audio output only**: `20 x $0.0005 = $0.01`. If the request also has 1,000 billed input tokens, add `$0.001`, giving an estimated total of **$0.011**. With the same 1,000 input tokens and six seconds of output, the estimate is `$0.001 + $0.003 = $0.004`. These are worked usage assumptions, not measured calls or a promise of clip length.

## Estimate, then reconcile one request

1. **Choose the exact model and text.** Record the published input and output rates and the expected request count.
2. **Estimate both input tokens and output duration.** Enter both into the calculator. For example, a narration budget needs longer audio assumptions than a short confirmation prompt.
3. **Request a sample with `response_format: "wav"`.** Compare the returned audio with the request's usage and charge in the console.
4. **Replace assumptions with observed quantities.** Add the text-input component and the audio-output component before multiplying by expected calls. Include repeated generations when estimating the workload.

A valid uncompressed WAV with a declared data length allows duration to be calculated from the audio data bytes divided by the byte rate. Do not divide the whole file size by that rate, because container headers are not audio samples. This is a way to check duration; it does not replace the billed input-token record.

The calculator uses current customer rates and excludes top-up fees and taxes. It is a planning estimate: ledger rounding and the proxy's token or duration estimates can produce a small difference from a hand calculation.

## Format and measurement limits

**WAV.** A valid header with a usable data length provides the most direct duration measurement. Malformed headers or streaming placeholders can limit that measurement, so a WAV extension alone is not proof of exact duration.

**Compressed formats.** The current proxy estimates non-WAV audio duration from the declared payload size and a nominal bitrate. Variable bitrate or container overhead can make that estimate differ from playback duration. Choose WAV when comparing billed audio duration with your own measurement.

**Unmeasurable audio.** When the proxy cannot establish a duration, it does not invent an audio-duration quantity. The input charge and any configured per-request charge can still apply; the current published TTS model has no flat per-request price. A missing duration is a record to investigate, not proof that the whole request was free.

**Implausible estimates.** An estimate above the proxy's plausibility limit is discarded as a duration measurement. This is not a guarantee that the request is rejected or that no input charge occurs.

## Budget the spoken result, including rejected takes

For a voice feature, keep the input text fixed and measure actual audio duration. A pronunciation correction may create a second paid take; include both in the cost of the accepted clip. The [Gemini pricing documentation](https://ai.google.dev/gemini-api/docs/pricing) explains its own units, while the KeepRouter model page sets the customer rate for this route. Avoid estimating all languages with one words-per-minute assumption: measure a representative sample at the selected voice and speed.

## Do not reuse the conversion across every speech model

Other speech products may use characters, tokens, requests, or different audio conversions. A familiar endpoint name does not establish the billing unit. Check the exact model's input and output prices before carrying this formula to another service; a missing price should remain unknown rather than become zero.

The [multimodal model API feature](/features/multimodal-models) explains the route-specific endpoints. The [video billing guide](/blog/choosing-a-video-model-tier) covers video duration pricing, which does not by itself describe the text-input component of a TTS request.


## Frequently asked questions

### Does the per-second rate include the input charge?

No. The current Gemini TTS route has a text-input rate and a separate audio-output rate. Add input tokens multiplied by their price to audio seconds multiplied by the per-second price.

### Why is my bill not proportional to character count?

Characters are not the complete billed quantity. KeepRouter estimates input tokens and measures or estimates generated audio duration; voice, punctuation and pacing can change the output component.

### Does WAV guarantee the exact bill before generation?

No. A valid WAV header can support direct duration measurement after generation, but the input charge still applies and malformed or placeholder lengths can limit measurement. The final record also follows ledger rounding.

### Is unmeasurable audio free?

Do not assume so. No duration quantity is added when measurement fails, but the input charge and any configured per-request charge can still apply. Inspect the request record.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [Google Cloud text-to-speech pricing (audio token rates and the tokens-per-second note)](https://cloud.google.com/text-to-speech/pricing)
2. [Google Gemini API pricing (speech models and audio token rates)](https://ai.google.dev/gemini-api/docs/pricing)
3. [KeepRouter current Gemini TTS customer prices](https://keeprouter.com/models/gemini-3.1-flash-tts-preview)

## Related guides

- [Multimodal model APIs](https://keeprouter.com/features/multimodal-models.md)
- [Async video generation API: submit, poll, and pay per second](https://keeprouter.com/blog/async-video-generation-api.md)
- [models](https://keeprouter.com/models.md)
- [errors](https://keeprouter.com/docs/errors.md)
- [gemini 3.1 flash tts preview](https://keeprouter.com/models/gemini-3.1-flash-tts-preview.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)

## Check this model's price and API

See the current customer rate, supported endpoint and setup example for the model discussed here.

[View model and pricing](https://keeprouter.com/models/gemini-3.1-flash-tts-preview)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
