# Async video generation API: submit, poll, and pay per second

> Video generation is a task queue, not a request/response call, and it is billed by the second of clip. Here is the contract, the polling rules, and the three failure cases that decide whether your queue double-charges.

_Published 2026-09-12 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 7 minute read_

![Asynchronous video task flow: a submit call enters a task canvas, returns a task id, and a dashed poll loop brings the result back](https://keeprouter.com/editorial/blog/async-video-generation-api.png)

_One submit call creates the task; the render completes asynchronously and is collected by polling._

**Short answer:** KeepRouter video generation uses an asynchronous task API: you submit work, receive a task id, and poll until the clip is ready. KeepRouter exposes that flow on `POST /v1/video/generations` and prices it by the **second of clip** rather than by tokens. Text-to-speech also meters audio output by duration, but its text-input charge is separate.

That has practical consequences for the code you write, and most of them show up after the first successful test call rather than during it.

## Why a synchronous call cannot work here

A five-second 1080p clip can occupy a GPU for a minute or more. If the submit call held the connection open, every client would need a timeout far beyond its normal request budget, and any proxy between you and the model would eventually cut a healthy render in half. Splitting the work into a submit and a poll keeps every individual HTTP call short and makes progress observable.

The cost of that split is state. A task id is a handle to work that is still running, so it has to survive a process restart, a redeploy, and a client that lost its in-memory map. Persist the task id next to whatever your application considers the unit of work, a job row, an order, a draft, and treat the poll as a resumable step rather than a loop that lives inside one function call.

## The contract, end to end

Submit the render. The response contains the task id and the initial state, and nothing else is needed to proceed:

```bash
curl https://keeprouter.com/v1/video/generations \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"wan2.7-t2v","prompt":"a paper plane gliding over a sunlit desk","duration":5}'
```

```json
{
  "id": "task_QeWgyFbrwjInjPMERqtdyFvehlMM4qgX",
  "task_id": "task_QeWgyFbrwjInjPMERqtdyFvehlMM4qgX",
  "object": "video",
  "status": "queued",
  "progress": 0,
  "created_at": 1789197093,
  "completed_at": null,
  "url": null,
  "error": null
}
```

Then poll. **Both `GET /v1/videos/{id}` and `GET /v1/video/generations/{id}` are the same operation**, because clients differ on which convention they already use; pick one and keep it. `status` moves through `queued` and `in_progress` to `completed`, `failed`, or `canceled`, and `progress` is a percentage when the provider reports one. `url` appears only once the clip exists, and it is a **signed link that expires**. Copy the file to your own storage when you first see it rather than handing the provider's URL to an end user.

Polling is a read. It is authenticated like every other call, it is not billed, and it does not count against your model budget. A five-second interval is a reasonable default; a tighter loop mostly adds requests.

**You do not have to poll at all.** Pass `webhook_url` on the submit call and the gateway POSTs a signed event when the task reaches a terminal state, carrying the same fields the poll returns. The delivery is signed with an HMAC over the timestamp and the raw body, so you can verify it before trusting it, and it is retried on a bounded schedule (about 1 minute, 15 minutes, 1 hour, then 6 hours) if your endpoint is down. Only public https destinations are accepted: internal, loopback and cloud-metadata addresses are refused before the task is ever submitted, and redirects are not followed.

Keep the poll as your fallback regardless. A webhook can be lost, and the poll response reports the delivery state (`webhook.status`, `attempts`, `last_error`), so a broken endpoint is visible rather than silent. Every poll response also carries `expires_at`, which is the moment the gateway stops tracking the task and the signed clip URL stops working, so a client can decide when it must fetch the file without computing the window itself.

## The billed quantity is the clip length

Video output is metered by **second of clip**. Text-to-speech audio output is metered by **second of generated audio**, with a separate text-input charge where published. Do not treat either duration-based output rate as a price per million text tokens. Model pages state the output unit (`per second of video`, `per second of audio`), and the machine-readable price feed carries an explicit `output_price_unit` column.

Where a provider charges more for a higher resolution, KeepRouter publishes **one model id per priced tier**: `wan2.7-t2v` is the 720p tier and `wan2.7-t2v-1080p` is the 1080p tier, each with its own rate. The gateway pins the tier it priced, so a request cannot be rendered above the tier you are paying for. If you ask for a resolution the id was not priced at, the call is refused with `resolution_not_priced` instead of being silently upgraded or downgraded. If you omit `duration`, a default is pinned and billed, so the clip you receive is the clip you paid for.

That last point is the one worth internalising: **the amount is known before the render runs**, because the provider charges at submission for the duration you asked for. Your own cost accounting can therefore be exact at submit time rather than reconciled after the fact.

## Failure handling that does not double-charge

Three situations need different responses, and treating them alike is how queues end up paying twice.

**The submit call times out.** You do not know whether the task was created. Do not blind-retry: a retry creates a second render and a second charge. Instead, make your submit idempotent at your own layer: one job row, one submit. If you must retry after a timeout, treat an unexpected extra task as the cost of that decision rather than the default path.

**The task fails.** Providers charge when the task is accepted and do not refund a task they later reject, and the gateway mirrors that rather than absorbing it. The poll tells you exactly what happened: `status` becomes `failed` and `error.message` carries the provider's own reason, for example a duration outside the range that model accepts. Read it before assuming the gateway dropped the work.

**The task id stops resolving.** A task is readable by the account that submitted it, so another account gets the same `404` as an unknown id. That is deliberate, because the poll response carries a signed download URL. KeepRouter also stops tracking a task after **24 hours**, matching how long the signed URL lives; after that the poll returns `task_expired`, and the `expires_at` field on every earlier poll response names that moment exactly. Store the clip, not the task id.

## Design the waiting screen around the task ID

Keep one application job linked to the returned task ID. If the browser reloads, resume polling that ID instead of submitting another render. Show queued, running, completed and failed as different states; elapsed time is not proof of progress. The [KeepRouter API reference](https://keeprouter.com/api/docs) defines the submit and status operations. Once the result is ready, save it according to the returned URL lifetime rather than treating a temporary download URL as permanent storage.

## What to check before shipping

Run one render of the exact duration and resolution you intend to use in production, and confirm four things: the task reaches `completed`, the clip downloads from the URL you were given, the charge equals duration multiplied by the published per-second rate for that tier, and a deliberately invalid request is refused before any upstream work happens. That last check is the one that tells you whether your validation runs early enough to protect the account.

The [model catalog](/models) is the current list of callable video ids and their per-second rates, and the [API reference](/docs/quickstart) covers the surrounding account and key setup. If you are comparing gateways on this class of endpoint, the [multimodal model API feature](/features/multimodal-models) explains why route-specific endpoints are published per model rather than flattened into one universal request, and the [error reference](/docs/errors) documents each code above.


## Frequently asked questions

### How long does a video render take?

A short clip typically finishes in under a minute, and longer or higher-resolution renders take more. Because the submit call returns immediately, the render time does not affect your request timeouts; poll every few seconds and persist the task id so a restart does not lose it.

### Am I charged if the task fails?

Yes. The provider charges when it accepts the task and does not refund work it later rejects, so the gateway mirrors that instead of absorbing the difference. Poll the task and read error.message to see the provider's own reason before retrying.

### Why are there two poll paths?

GET /v1/videos/{id} and GET /v1/video/generations/{id} are the same operation. Video APIs in the wild use both conventions, so the gateway accepts either rather than making you adapt to one.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [Google Cloud text-to-speech pricing (audio is billed per token, with a documented tokens-per-second rate)](https://cloud.google.com/text-to-speech/pricing)
2. [Google Gemini API pricing (duration-metered video and audio models)](https://ai.google.dev/gemini-api/docs/pricing)
3. [OpenAI API pricing (audio token rates for speech models)](https://platform.openai.com/docs/pricing)
4. [KeepRouter API reference](https://keeprouter.com/api/docs)

## Related guides

- [Multimodal model APIs](https://keeprouter.com/features/multimodal-models.md)
- [AI gateway guide: what it controls, when to use one, and how to adopt it](https://keeprouter.com/blog/ai-gateway-guide.md)
- [LLM failover design guide: recover requests without hiding unsafe retries](https://keeprouter.com/blog/llm-failover-design-guide.md)
- [errors](https://keeprouter.com/docs/errors.md)
- [wan2.7 t2v](https://keeprouter.com/models/wan2.7-t2v.md)

## Check the behavior your application needs

Use the documented request format and controls when applying the example to your application.

[Read the product reference](https://keeprouter.com/api/docs)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
