Choosing a video model tier: cost per second, resolution, and input type
Video tiers in the catalog differ by input type, resolution, and price per second. This is how to pick one without guessing: the cost math, the minimum durations, and where each family fits.
Published 2026-09-12 · Updated 2026-09-29 · KeepRouter Editorial · 6 minute read

Short answer: the catalog publishes one id per priced tier, so choosing is a decision about three things you already know before you write any code: what you are feeding the model, how long the clip needs to be, and how much resolution is worth to you. There is no ranking to consult, because this catalog does not publish benchmark scores.
Start with the input type, not the price
The families in the catalog accept different inputs, and that usually decides the choice before cost enters the picture.
| What you have | What you need | Shapes available |
|---|---|---|
| A written description | A clip from scratch | wan2.7-t2v, happyhorse-1.0-t2v, veo-3.1-lite, veo-3.1-fast, veo-3.1, and the PixVerse tiers |
| A still image to animate | Motion that follows that frame | wan2.7-i2v and happyhorse-1.0-i2v |
| One or more reference images | A clip that follows the references | wan2.7-r2v and happyhorse-1.0-r2v |
| An existing clip to change | An edited or restyled clip | wan2.7-videoedit and happyhorse-1.0-video-edit |
The HappyHorse input-driven ids take their input through a media array rather than a prompt alone: [{"type":"first_frame","url":"https://..."}] for image-to-video, reference_image for reference-to-video, and video for video editing, where the input clip must be at least three seconds long. The URL has to be fetchable by the provider, so a link behind your own authentication will not work, and a wrong type costs nothing: the model refuses the request and names the values it accepts before any render is billed.
If you have a still and want motion, a text-to-video model is the wrong tool even when it is cheaper per second, because it will not preserve your frame. Decide the family first, then compare the tiers inside it.
The cost math, with the real numbers
Video is billed by the second of clip, so the arithmetic is plain multiplication: cost = seconds x published rate for that id. The rate for every id is on its model page and in the catalog. Using published rates, a five-second clip costs:
| Id | Tier | Rate per second | 5-second clip |
|---|---|---|---|
pixverse-v6-360p | 360p | $0.025 | $0.125 |
pixverse-v5.6-360p | 360p | $0.035 | $0.175 |
pixverse-v3.5-360p, v4, v4.5, v5, v5.5 | 360p | $0.045 | $0.225 |
wan2.7-videoedit | 720p | $0.086012 | $0.430 |
minimax-h3-768p | 768P | $0.09 | $0.450 |
wan2.7-t2v, wan2.7-i2v, wan2.7-r2v | 720p | $0.10 | $0.500 |
veo-3.1-lite | 720p | $0.05 | $0.250 |
veo-3.1-lite-1080p | 1080p | $0.08 | $0.400 |
minimax-h3-1080p | 1080P | $0.13 | $0.650 |
happyhorse-1.0-t2v, happyhorse-1.0-i2v, happyhorse-1.0-r2v, happyhorse-1.0-video-edit | 720p | $0.14 | $0.700 |
wan2.7-t2v-1080p, wan2.7-i2v-1080p, wan2.7-r2v-1080p | 1080p | $0.15 | $0.750 |
wan2.7-videoedit-1080p | 1080p | $0.143353 | $0.717 |
happyhorse-1.0-t2v-1080p, happyhorse-1.0-i2v-1080p, happyhorse-1.0-r2v-1080p, happyhorse-1.0-video-edit-1080p | 1080p | $0.24 | $1.200 |
veo-3.1-fast | 720p | $0.10 | $0.500 |
veo-3.1-fast-1080p | 1080p | $0.12 | $0.600 |
veo-3.1-fast-4k | 4K | $0.30 | $1.500 |
veo-3.1, veo-3.1-1080p | 720p / 1080p | $0.40 | $2.000 |
veo-3.1-4k | 4K | $0.60 | $3.000 |
Two things are worth reading off that table. First, the spread between the cheapest and the most expensive 5-second clip is now more than twenty times, so a tier choice made carelessly is a real budget line. Second, the jump from a base tier to its high-resolution sibling is not uniform: for the Wan 2.7 text-to-video pair it is 1.5x, for HappyHorse it is about 1.7x, for MiniMax it is about 1.4x, for Veo 3.1 Lite it is 1.6x, for Veo 3.1 Fast it is 1.2x, and for the flagship Veo tier the two resolutions cost exactly the same. The 4K step is a separate jump again: 2.5x the 1080p rate on Veo 3.1 Fast and 1.5x on the flagship. Comparing "the 1080p one" across families without checking the pair is how estimates drift.
Do the expensive part last. Prototype at the cheapest tier that accepts your input, settle the prompt and the duration, and only then re-run the final takes at the resolution you intend to ship. A prompt iteration that costs $0.125 per attempt is a different workflow from one that costs $0.75.
Duration is a billed input, and minimums differ
duration is not a hint: it is the quantity you are charged for, and the gateway sends the value it bills, so the clip you receive matches the amount you paid. If you omit it, a default is pinned and charged rather than left to the provider.
Minimum durations differ by family. MiniMax H3 rejected a 2-second request with a parameter error and accepted the same request at 6 seconds, so treat short durations as something to verify per model rather than assume. Providers also cap the maximum, and a value outside a model's accepted range surfaces as a failed task rather than a rejected submit, which is why the request log and the task's own error message are worth reading before you retry.
What resolution actually buys
A higher tier charges more because the provider encodes more pixels per frame, not because the model changed. Whether that matters depends on where the clip is displayed: a full-width hero on a desktop page and a small inline preview have very different requirements, and the same 360p file that looks soft in one looks fine in the other. Because each tier is a separate id with its own published rate, you can measure that difference on your own content and decide with evidence instead of defaulting to the highest tier.
A five-second calculation is not a supported-duration promise
The table illustrates rate multiplication; it does not establish that each model accepts a five-second request. Check the permitted duration and input type on the selected model page. For a useful trial, render the same permitted brief at the available tiers and judge identity preservation, motion, legibility and framing. Count rejected renders in cost per usable clip. Historical price examples should be replaced with the current catalog rate before budgeting.
A short selection sequence
- Filter by input type. Drop every id whose family does not accept what you have.
- Set a duration budget. Multiply the candidate rates by the clip length you actually need, including retries.
- Pick the cheapest tier that satisfies the display size. Render one sample at that tier and one at the next tier up, then compare them where they will be shown.
- Verify the exact request once. Confirm the task reaches
completed, the clip downloads, and the charge equals duration multiplied by the published rate for that id.
The model catalog carries the current rates, and how the asynchronous flow works covers the submit-and-poll mechanics that every one of these ids shares. If you are still deciding whether your use case belongs behind a gateway at all, the multimodal model API feature explains why each modality keeps its own endpoint instead of sharing one universal request.
Frequently asked questions
Does a higher price mean a better clip?
No. The published rate follows what the provider charges, which tracks resolution, duration, and the model family, not a quality ranking. This catalog does not publish benchmark scores, so treat the price as a cost input and evaluate output on your own material.
Why can the same prompt cost different amounts on different ids?
Each id pins a specific tier. wan2.7-t2v is the 720p tier and wan2.7-t2v-1080p is the 1080p tier, so the same prompt and duration produce different charges on purpose. If you send a resolution the id was not priced at, the request is refused rather than rendered at a tier you did not pay for.
Is there a 4K or audio tier?
Yes to 4K: `veo-3.1-fast-4k` and `veo-3.1-4k` are published at $0.30 and $0.60 per second. There is no separate audio tier to choose, because a default render comes back with synced audio and that is the rate these ids bill. Tiers stay out of the catalog when their exact per-second cost cannot be established from the provider at all, rather than being published at a guessed rate.
Sources reviewed
Article last reviewed 2026-09-29
- [1] Google Gemini API video documentation (video generation parameters)
- [2] Google Gemini API pricing (duration-metered video models)
- [3] model page