Webhooks for long-running AI jobs: signed, retried, never the only path

A webhook removes the need to poll in the common case, not the need to poll at all. The contract, the retry ladder, signature verification, and the destination rules that keep a callback from becoming an SSRF primitive.

Published 2026-09-13 · Updated 2026-09-29 · KeepRouter Editorial · 8 minute read

Push notification diagram: a submit block, an arrow to a receiver block, a confirmation node with a signature line, and a retry ladder of five shrinking bars
A signed delivery announces the result once; the retry ladder decides how long a broken receiver can stay broken.

Short answer: a webhook does not replace the poll, it removes the need for one in the common case. You pass webhook_url when you submit a long-running job such as a video render, and the gateway POSTs a signed event when the task reaches a terminal state. The poll remains the source of truth, it is unbilled, and it reports the delivery state of the webhook, so a receiver outage stays visible while the result is still inside its 24-hour retrieval window.

That framing matters because most webhook integrations are designed as if the callback were guaranteed. It is not. A callback can be dropped by a restart, rejected by a deploy, or land while your service is rate limited. Designing for "the webhook is a hint, the poll is the record" is the difference between an integration that recovers on its own and one that pages somebody at 03:00.

Why long jobs ask for push

A five-second 1080p clip can occupy a GPU for a minute or more, so the video API is a task queue: submit, receive a task id, then learn the outcome later. The default way to learn it is polling, and polling is fine at small scale. It stops being fine when the client is a serverless function, because every poll is an invocation, or when the client is a browser tab that may be closed, or when the job takes longer than any interval you would call polite.

A callback inverts the direction: instead of asking "is it done yet" on a timer, your service is told once, with the result attached.

What you havePollingWebhook
A short job under a few secondsSimplerOverkill
A serverless worker that must wake on completionCosts an invocation per pollOne delivery per task
A browser or mobile client that may be closedProgress is lost with the tabA server receives completion independently
A pipeline step that blocks on the resultPoll loop inside the stepThe step ends, the event restarts it

The delivery contract

Submit as usual and add webhook_url:

curl https://keeprouter.com/v1/video/generations \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"veo-3.1-fast","prompt":"a paper plane over a sunlit desk","duration":4,
       "webhook_url":"https://hooks.example.com/keeprouter"}'

The submit response carries the task id and a one-time signing secret. Keep that secret: it is returned once, in that response, and the poll never returns it again. When the render finishes, the gateway POSTs an event whose body contains the same fields the poll would have returned:

{
  "id": "task_QeWgyFbrwjInjPMERqtdyFvehlMM4qgX:completed",
  "type": "video.task.completed",
  "created_at": "2026-09-13T01:00:00.000Z",
  "data": {
    "id": "task_QeWgyFbrwjInjPMERqtdyFvehlMM4qgX",
    "object": "video",
    "model": "veo-3.1-fast",
    "status": "completed",
    "seconds": 4,
    "url": "https://.../renders/xxx.mp4?Expires=...&Signature=...",
    "error": null,
    "tracked_until": "2026-09-14T01:00:00.000Z"
  }
}

Three headers carry the delivery metadata: kr-event-id (stable across retries), kr-attempt (which attempt this is) and kr-signature.

A 2xx response means delivered. The retry policy is deliberately small and slow, and it is worth memorising because it decides how long your endpoint can be broken without losing an event:

AttemptWhenTrigger for the next one
1At the moment the task finishes, in the same request that observed itResponse was not 2xx
2About 1 minute later408, 425, 429 or any 5xx
3About 15 minutes after thatsame
4About 1 hour after thatsame
5About 6 hours after that, then it stopssame

Any other 4xx is treated as permanent, because a receiver that answers 400 to a well-formed event will answer 400 again, and 3xx is permanent by design: redirects are not followed, which is a security property rather than an oversight (see below).

Verifying a delivery

The signature is HMAC-SHA256 over <timestamp>.<raw body>, hex encoded, in a header shaped t=1730000000,v1=<hex>. The timestamp is inside the signed material so a captured delivery cannot be replayed with a fresh header. In Node:

import crypto from "node:crypto";

export function verify(rawBody, header, secret, toleranceSeconds = 300) {
  if (typeof header !== "string") return false;
  const parts = Object.fromEntries(header.split(",").map((p) => p.split("=")));
  const age = Math.abs(Math.floor(Date.now() / 1000) - Number(parts.t));
  if (!Number.isFinite(age) || age > toleranceSeconds) return false;
  const expected = crypto
    .createHmac("sha256", secret)
    .update(`${parts.t}.${rawBody}`)
    .digest("hex");
  const received = parts.v1;
  if (typeof received !== "string" || !/^[0-9a-f]{64}$/i.test(received)) return false;
  return crypto.timingSafeEqual(Buffer.from(expected, "hex"), Buffer.from(received, "hex"));
}

Three details decide whether this holds up in production. Reject a missing or malformed signature before the constant-time comparison. Verify against the raw body, not a re-serialised object, because key order and whitespace are part of what was signed. And compare with a constant-time function, not ===.

Idempotency, on both sides

A retry posts byte-identical content: same event id, same body, same signature timestamp. That is intentional, because it lets you deduplicate on kr-event-id with a unique index and stop thinking about it:

INSERT INTO webhook_events (event_id, received_at)
VALUES (?, now()) ON CONFLICT (event_id) DO NOTHING;

If the insert reports no row, you have already handled this event. The alternative, deduplicating on the clip URL or on a timestamp, breaks the moment two tasks finish in the same second.

Keep the poll wired up

Every poll response now carries the delivery state, alongside the moment the tracking window ends:

{
  "id": "task_...",
  "status": "completed",
  "url": "https://.../renders/xxx.mp4?Expires=...",
  "expires_at": "2026-09-14T01:00:00.000Z",
  "webhook": { "status": "delivered", "attempts": 1, "last_error": null, "delivered_at": "2026-09-13T01:00:05.000Z" }
}

That is what makes a webhook safe to rely on. If webhook.status is pending with a last_error, your endpoint is the problem and you can see it without a support ticket. If it is failed, the event will not be retried again and the poll is the only way left to collect the clip. expires_at tells you the deadline for that: the gateway tracks a task for 24 hours, matching the provider's signed download window, and after that the clip URL is gone too.

Why the destination rules are strict

The gateway is the one making the request, from inside Cloudflare, with its own database and storage bindings in scope. A destination it will POST to is therefore a security boundary, not a convenience. Only https on port 443 with a public DNS hostname is accepted; internal names, loopback addresses, private ranges and cloud metadata addresses are refused before the task is even submitted, and redirects are never followed, because a 302 to a metadata endpoint would walk straight past every rule above.

A rejected destination costs nothing and creates no task. The request fails with invalid_webhook_url and names the reason, and it is worth deciding the destination before you write the submit call rather than after.

When not to use a webhook

Jobs that finish in a second. A callback adds a public endpoint, a secret to rotate and a retry queue to monitor. For a request that returns inline, that is a poor trade.

A receiver that cannot verify signatures. An unverified webhook endpoint is worse than polling: it is a public URL that mutates your state on request, and anybody who learns the URL can drive it. If you cannot verify, poll.

A team with no durable endpoint. A webhook needs somewhere stable to land. A developer laptop that sleeps is not that, and the retry ladder will exhaust itself politely while nothing is listening.

Test duplicate delivery with an inert job

Send the same signed test event to a staging handler twice. The second delivery should acknowledge receipt without repeating the business action. Then process a status poll after the callback and confirm it does not create a second completion. Stripe’s webhook guide documents duplicate-event handling as a general integration concern; KeepRouter’s event fields and signature remain defined by its own API reference. Do not copy Stripe header names into this handler.

The pre-ship checklist

  1. Verify the signature against the raw body, in constant time, and reject old timestamps.
  2. Deduplicate on kr-event-id with a unique constraint in your own storage.
  3. Return 2xx fast, then process asynchronously. A receiver that does the work before responding will time out at 5 seconds on a slow day and earn a retry.
  4. Watch webhook.status on the poll you were already making, or alert on failed and pending with attempts above two.
  5. Keep the poll path implemented, because it is the only route to a clip whose event was never delivered.

The async video flow explains the submit and poll mechanics this builds on, the video generation reference documents every field, and model pages state the per-second rate a finished render is billed at. If your jobs are not video, the same contract applies to any route that takes a webhook_url, and the error reference covers invalid_webhook_url.

Frequently asked questions

Do I still need to poll if I use a webhook?

Yes, keep the poll as a fallback. A webhook can be lost while your endpoint is down, and the retry ladder stops after five attempts over about seven and a half hours. The poll response reports the delivery state, so you can see a failure and collect the clip yourself.

How do I verify a delivery really came from the gateway?

Recompute HMAC-SHA256 over the string formed by the timestamp, a dot, and the raw request body, using the signing secret returned once at submit, then compare it with v1 in the kr-signature header in constant time. Reject timestamps older than a few minutes so a captured delivery cannot be replayed.

Are webhook deliveries billed?

No. Deliveries are not billed, and neither is polling. Only the task itself is billed, at its published per-second rate multiplied by the clip length you requested.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Stripe webhooks documentation (retries, idempotency and signature verification)
  2. [2] GitHub webhooks documentation (delivery ids, redelivery and validation)
  3. [3] OWASP Server-Side Request Forgery Prevention Cheat Sheet (destination validation)

Related guides

← All posts · Models & pricing · Get an API key