Multimodal embeddings: one vector space for text, images, and video

Some embedding models accept images and video, not just text, and they are rejected on the standard embeddings route. Here is the shape to send, how usage is counted, and what it replaces.

Published 2026-09-12 · Updated 2026-09-29 · KeepRouter Editorial · 4 minute read

Three typed inputs (text, image, video) converging into one vector block
Typed text and media inputs form a vector representation; define the item boundary before building a retrieval index.

Short answer: an embedding is a vector, and a multimodal embedder produces that vector from text, images, and video in the same space. That single property is the reason to use one: an image and a sentence about that image land near each other, so retrieval across both becomes one query against one index instead of a pipeline that transcribes, captions, and then embeds the captions.

The endpoint is different, and that is not a quirk

Some providers serve these models on a separate operation from ordinary text embeddings. Sending such a model to the standard embeddings route returns an explicit error saying the model does not support that API. That is why this catalog publishes them under their own modality and points their pages at POST /v1/embeddings/multimodal rather than at the text endpoint.

Publishing the wrong endpoint would be worse than publishing none: a client could read the model page, copy a working-looking example, and get a rejection. The rule this catalog follows is that the endpoint on a model page is the endpoint that actually serves that model.

The request shape

Input is an array of typed parts rather than a bare string, so a text-only request still works but has to say so:

curl https://keeprouter.com/v1/embeddings/multimodal \
  -H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"doubao-embedding-vision-251215","input":[
        {"type":"text","text":"a paper plane on a desk"},
        {"type":"image_url","image_url":{"url":"https://keeprouter.com/logo-512.png"}}
      ]}'

The gateway normalises the response to the OpenAI list form, so existing code that reads data[0].embedding keeps working even though the provider returns a single object. It also accepts a plain string or an array of strings and wraps them into parts, which means a text-only call does not need a different code path.

How usage is counted

The response reports usage that separates media from text. A request carrying one small image and one short sentence came back as 362 prompt tokens, of which 334 were image tokens and 28 were text tokens. That breakdown is the honest way to see where the cost comes from, and it explains why a request that looks small can still be counted in the hundreds.

The provider charges more for image tokens than for text tokens. This catalog's pricing table carries one input rate per model, so the published rate is the text rate and media tokens are billed at it. In practice that is a discount the operator absorbs rather than an overcharge, and at these rates the whole difference is a fraction of a cent per call. If image-heavy traffic ever became material, the fix is a second input rate, not a higher text rate.

When this is the right tool

Use a multimodal embedder when the query and the corpus are of different kinds and you want one ranking across both:

  • Visual search inside a product catalog, where a shopper's sentence has to match a photo.
  • Deduplication across a mixed archive of screenshots, frames, and text notes.
  • Retrieval for multimodal assistants, so a question can pull a relevant frame without a captioning step in front of it.

Reach for something else when the task needs words rather than similarity. If you must read the text in an image, that is OCR. If you must describe what is in a picture in sentences, that is a vision model. An embedder returns numbers, and its job is to make similar things close, not to explain them.

Define one catalog item before creating vectors

In a product search prototype, a product image and its caption can describe one item; unrelated images should remain separate items. Keep the same representation recipe for the index and queries, then check retrieval on known image-text pairs. The KeepRouter API reference defines this route; the Doubao retrieval tutorial adds a finite-vector check and evaluation method. A successful embedding response does not show that your search results are useful.

Practical notes

Dimensions are fixed by the model. The model here returns a 2048-dimension vector. Store that width in your schema and pin it, because switching embedders later means re-embedding the corpus rather than changing a query.

Context is large but not unlimited. The published context window is 131,072 tokens, which is generous for documents and video-derived frames, and still finite. Long inputs are truncated or rejected by the provider depending on the operation, so chunk deliberately.

Embed once, then reuse. Embeddings only pay off when the corpus is embedded once and queried many times. Treat the vector store as the durable artefact and the embedding call as a build step, not something a request handler does inline.

The multimodal model API feature explains the per-modality endpoint rule, the catalog lists current rates and the exact endpoint for each model, and the video task guide covers the other asynchronous path in this catalog.

Frequently asked questions

Why is this not just /v1/embeddings with an image field?

Because the providers that offer it implement it as a separate operation. The model is rejected on the standard embeddings route with an explicit error, so publishing it there would advertise an endpoint that cannot serve it. The gateway gives these models their own modality and points their pages at the multimodal path.

Are images billed at the same rate as text?

The provider charges more for image tokens than for text tokens, but the pricing table carries a single input rate, so the published rate is the text one and media tokens are billed as text. That is a discount the operator absorbs, never an overcharge, and it is worth fractions of a cent per call at these rates.

What inputs are accepted?

Text, images, and video, as typed parts in one request. The model in this catalog reports a 2048-dimension vector and a 131,072-token context window, and its usage breakdown separates image tokens from text tokens so you can see how a request was counted.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Google Gemini API embeddings documentation (multimodal input, per-type limits)
  2. [2] OpenAI embeddings guide (request and response shape)
  3. [3] KeepRouter API reference

Related guides

← All posts · Models & pricing · Get an API key