Image-to-video and video editing APIs: the media contract

Animating a still or editing a clip uses the same task API as text-to-video, with the input in a media array whose type differs per model. The shapes, the rules the provider enforces, and what each tier costs per second.

Published 2026-09-13 · Updated 2026-09-29 · KeepRouter Editorial · 7 minute read

Input-driven video diagram: a still frame on the left, an arrow, then a filmstrip of three output frames with the middle one highlighted
A first frame is animated rather than replaced; the media array is what tells the model which role the input plays.

Short answer: image-to-video, reference-to-video and video editing use the same task API as text-to-video. What changes is where the input goes: a media array of { "type": "...", "url": "..." } objects, whose type differs per model. The billed quantity is unchanged, it is still duration multiplied by the published per-second rate for that tier.

That one array is where most first attempts fail, and the failures are informative rather than mysterious: the model refuses the request and names the values it accepts. Knowing the shape in advance turns a debugging session into a single request.

The four input shapes

A video model is chosen by what you already have, not by price. The catalog publishes ids for each shape:

What you haveWhat the model doesPublished ids
A written descriptionGenerates a clip from scratchwan2.7-t2v, happyhorse-1.0-t2v, veo-3.1-lite, veo-3.1-fast, veo-3.1, PixVerse tiers
A still imageAnimates that frame, keeping it as the first framewan2.7-i2v, happyhorse-1.0-i2v
One or more reference imagesGenerates a clip that follows the referencewan2.7-r2v, happyhorse-1.0-r2v
An existing clipEdits or restyles that clipwan2.7-videoedit, happyhorse-1.0-video-edit

If you have a still and the render must preserve it, a text-to-video model is the wrong tool even at a lower rate: nothing in the request tells it what your frame looks like.

The media array, per model

The input travels in a top-level media array whose items name their own role:

Model familymedia[].typeurl points at
happyhorse-1.0-i2vfirst_framethe still to animate
happyhorse-1.0-r2vreference_imagethe image to follow
happyhorse-1.0-video-editvideothe clip to edit, at least 3 seconds
wan2.7-i2vfirst-frame image fieldthe still to animate
wan2.7-videoeditsource clip fieldthe clip to edit

A concrete submit for image-to-video:

curl https://keeprouter.com/v1/video/generations \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"happyhorse-1.0-i2v","prompt":"slow cinematic push-in on the subject","duration":4,
       "media":[{"type":"first_frame","url":"https://cdn.example.com/frame.png"}]}'

And for video editing, with a clip long enough to be accepted:

curl https://keeprouter.com/v1/video/generations \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"happyhorse-1.0-video-edit","prompt":"restyle as a watercolour painting","duration":3,
       "media":[{"type":"video","url":"https://cdn.example.com/source.mp4"}]}'

Both return a task id, and both are collected from GET /v1/videos/{id} exactly like a text-to-video render. Nothing about the asynchronous flow changes.

Three rules the provider enforces

The URL has to be fetchable by the provider, not by you. A signed link that works in your browser, or a URL behind your own authentication, will fail: the model downloads the input itself. In testing this, a Wikimedia-hosted image and a well-known public sample bucket were both refused, while a file served from a public CDN worked. Host the input somewhere reachable, and prefer a stable URL over one that expires in minutes.

A video input has a minimum length. The editor rejected a 2.04-second clip with "duration should be at least 3s", so trim or extend your source before submitting. The error is explicit about the measured length, which makes it a one-edit fix.

A wrong type is free. The model answers with the values it accepts ("Input should be 'first_frame'", or a list when it takes more than one), before any render runs, and a request refused at that point is not billed. This is the cheapest possible way to confirm the contract for a model whose page you have not read yet.

What it costs

Per-second rates are published per id. Some input-driven tiers match their text-to-video sibling, while others do not, so use the exact model id instead of inferring a rate from the family name:

IdTierRate per second5-second clip
wan2.7-videoedit720p$0.086012$0.430
wan2.7-i2v, wan2.7-r2v720p$0.10$0.500
happyhorse-1.0-i2v, happyhorse-1.0-r2v, happyhorse-1.0-video-edit720p$0.14$0.700
wan2.7-i2v-1080p, wan2.7-r2v-1080p1080p$0.15$0.750
wan2.7-videoedit-1080p1080p$0.143353$0.717
happyhorse-1.0-i2v-1080p, happyhorse-1.0-r2v-1080p, happyhorse-1.0-video-edit-1080p1080p$0.24$1.200

The billed quantity is the clip length you ask for, not the length of the input. A 30-second source clip edited into a 3-second output is billed as 3 seconds.

Accepted durations are per model and enforced before the request leaves the gateway: the HappyHorse family takes 3 to 15 seconds, the Wan 2.7 family 2 to 15. Omit duration and the gateway sends the smallest length the model accepts, so an omitted field can never produce a request the provider refuses. Sending a length the model does not accept returns 400 invalid_duration naming the range, before any upstream call.

Getting a usable result

Choose the first frame like a shot, not like a photo. The model animates what is in the frame, including a busy background or a subject facing away from the light. A frame with a clear subject and a simple background gives the motion model somewhere to go.

Say what should move, and how. "Slow push-in", "hair moving in the wind", "the camera orbits the object" are instructions the model can act on. Leaving the prompt empty and hoping the first frame carries the shot wastes a render you already paid for.

Reference images and first frames are different jobs. A first frame pins the opening frame of the output. A reference guides style or subject across the clip without being copied frame for frame. If you need the exact input frame at the start, that is first_frame.

Prototype at 720p, ship at the resolution you actually display. For the HappyHorse rows above, the 1080p tier costs around 1.7x the 720p tier, so a prompt iteration at 720p is a different budget from one at 1080p.

Keep the source clip reasonable. For editing, a longer input does not cost more, but it does take longer to process, and only the requested output length is billed.

Make the input URL part of the test

Before rendering, verify that the exact input URL can be fetched without an interactive login and will remain valid long enough for the job. Use a non-sensitive test asset you have permission to process. A URL that opens in your signed-in browser may fail for the generation service. Follow the media fields in the KeepRouter API reference, then check the resulting first frame and subject identity. Successful submission alone does not establish that the intended image was used correctly.

Verify once, then automate

Run one render of the exact shape you intend to use and confirm four things: the task reaches completed, the clip downloads, the charge equals duration multiplied by the published rate, and a deliberately wrong media[].type is refused without a charge. Then persist the task id next to your own work item and either poll or subscribe a webhook for the completion.

The tier selection guide covers how to choose among the tiers once the input shape has narrowed the field, the async flow post covers the submit and poll mechanics, and the video generation reference lists every field. Model pages such as happyhorse-1.0-i2v state the rate, the accepted clip lengths and a runnable example with the right media type already filled in.

Frequently asked questions

Why did my image-to-video request fail with a missing input field?

Because the input has to travel in a media array, not in a prompt. Send media as a top-level array whose items are objects with type and url, for example type first_frame for image-to-video. A wrong or missing type is refused before any render runs, and is not billed.

What URL can I use for the input image or video?

One the provider can download: a public https URL that does not require your credentials and does not expire in minutes. Signed links that only work in your browser, and hosts that refuse automated downloads, both fail. A file on a public CDN or your own public bucket is the reliable choice.

Is image-to-video more expensive than text-to-video?

No. Input-driven tiers are published at the same per-second rates as their text-to-video siblings in the same family, so animating a still costs the same as generating from scratch at that tier. You are billed for the output duration you request, not for the length of the input.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Google Gemini API video documentation (image input and generation parameters)
  2. [2] Google Gemini API pricing (duration-metered video tiers)
  3. [3] KeepRouter API reference

Related guides

← All posts · Models & pricing · Get an API key