# Embedding model migration: reindex without losing retrieval quality

> Changing an embedding model changes the retrieval contract, even when the vector dimensions happen to match. Re-embed documents and queries into a new version, compare a fixed set of relevant results, then switch traffic with the old index still available for rollback.

_Published 2026-09-28 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Three typed inputs (text, image, video) converging into one vector block](https://keeprouter.com/editorial/blog/multimodal-embeddings-guide.png)

_Treat each embedding model and index version as one retrieval contract; backfill and evaluate the replacement before cutover._

**Short answer:** do not switch only the embedding model ID in a live RAG application. Keep the old document and query vectors paired while you build a second vector space from the same source documents. Evaluate retrieval and answer support on both paths, cut over a small cohort, and retain the old index until the new one is stable. Qdrant's [official migration tutorial](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/) documents both a parallel-collection and a named-vector route.

This is a migration method, not a report of a KeepRouter customer's outcome. The example counts below are hypothetical and the current model price must be read from the [live catalog](/models).

## Why a same-size vector is still a different space

An embedding is produced by a particular model, revision, input transformation and dimensionality setting. A database collection also has a distance metric and filters. Two models can emit vectors of the same length while arranging meanings differently. Comparing new query vectors to old document vectors can silently return plausible but wrong passages. OpenAI's [embedding guide](https://developers.openai.com/api/docs/guides/embeddings) also documents a dimensions option for some models; shortening a vector is an explicit model configuration, not a license to mix it with an old index.

Write a retrieval contract before changing traffic:

| Field | Old path | New path | Gate |
| --- | --- | --- | --- |
| Model and revision | Exact model ID | Exact target ID | Both documented and callable |
| Input preparation | Normalization and chunker version | Same or deliberately revised | Saved source can reproduce both |
| Vector schema | Dimension and distance | Target dimension and distance | Collection accepts target vectors |
| Permission filters | Tenant and document ACL | Same effective policy | Zero unauthorized matches |
| Answer stage | Fixed generator and prompt | Keep fixed for retrieval test | Difference attributable to retrieval |

## The migration sequence

1. **Freeze the evidence set.** Save at least one query for every high-value intent, known relevant document IDs, known no-answer queries and unauthorized-document probes. Version the labels. Do not label a query by asking only the model being replaced.
2. **Recover the source.** Confirm you can reproduce document text, media references, chunk IDs, deletions and access-control metadata. Old vectors alone are not sufficient to build new embeddings.
3. **Create a second vector space.** Qdrant describes either a blue-green collection or an additional named vector when the collection supports it. Record the new dimension and distance choice. Leave the old route serving.
4. **Backfill and dual-write.** Re-embed existing source records in resumable batches; send new and changed documents to both paths. Treat deletes and ACL changes as first-class events so a stale point cannot reappear after cutover.
5. **Compare offline.** Run the same queries against old and new paths. Inspect required-passage recall, rank of supporting passages, permission leakage, p95 retrieval latency and answer support with the answer model held fixed.
6. **Cut over and observe.** Route a small cohort to the new index and record index version with each case. Retain the old index, query model configuration and reversal switch until production outcomes match the acceptance gates.

Qdrant's tutorial warns that blue-green dual writes need special care for deletes and partial updates. Its named-vector route can leave the old vector in place during a transition, but availability depends on the collection shape and version. Choose the path your actual database supports; the migration principle is broader than one vendor's implementation.

## A small cost worksheet, with no invented price quote

Suppose, only to size work, a corpus has 10,000 source documents, four chunks per document, and an average of 500 billable input tokens per chunk. The backfill is about 10,000 × 4 × 500 = 20,000,000 embedding input tokens, before retries, media processing and changed documents. At a hypothetical rate R dollars per million input tokens, the model part is 20 × R. Add vector storage, dual-write duration, query traffic, evaluation calls, generation and operator time. A short-lived parallel index can cost more than the embedding call itself.

The [KeepRouter model page](/models/doubao-embedding-vision-251215) lists the exact route and live customer price for that catalog model. Some multimodal embedders use /v1/embeddings/multimodal, not /v1/embeddings; the [modality guide](/blog/multimodal-embeddings-guide) explains the request shape. Check whether the model and input type you require are actually published before putting them on the migration plan. KeepRouter does not run the vector store or promise that all provider-specific embedding parameters carry through.

## Define a rollback signal from retrieval outcomes

Before switching, choose fixed queries whose relevant documents are known. Compare recall at your chosen k, unsupported answers and permission filtering on both indexes. Keep query embeddings paired with the matching document index throughout the test. The [Qdrant embedding-migration tutorial](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/) explains a dual-representation migration approach. A complete vector count alone does not prove the new index should receive users.

## Decide with outcomes, not only vector counts

Passing a backfill count proves completion, not relevance. Require a reviewed query set and compare recall at the same top-k. A zero-result query should remain unanswered rather than becoming a confident unsupported answer. Permission probes must be strict: one leaked passage is a failure even if average recall improves. Keep answer generation fixed while comparing the retrievers, then evaluate a new generator as a separate change. The [retrieval team guide](/built-for/retrieval-applications) maps gateway records to application outcomes.

After cutover, report the exact index and embedding model version in your own trace. The gateway's request status, token use and charge explain the API call, while your retrieval trace explains which passages were selected. You need both to diagnose an answer regression.

## Frequently asked questions

### Can I reuse old document vectors with a new query model?

Do not assume so, even if dimensions match. Re-embed and evaluate a separate vector space built from the same source documents.

### Do I need a second collection?

Not always. Qdrant documents both blue-green collections and named vectors, subject to the existing collection and product version.

### What should block cutover?

Any permission leak, missing required passage, unsupported answer or unacceptable latency or cost regression against the agreed test set.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [Qdrant embedding model migration](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/)
2. [OpenAI embeddings guide](https://developers.openai.com/api/docs/guides/embeddings)
3. [KeepRouter live model catalog](https://keeprouter.com/models)
4. [KeepRouter OpenAPI](https://keeprouter.com/api/openapi.json)

## Related guides

- [Retrieval and RAG applications](https://keeprouter.com/built-for/retrieval-applications.md)
- [Multimodal embeddings: one vector space for text, images, and video](https://keeprouter.com/blog/multimodal-embeddings-guide.md)
- [doubao embedding vision 251215](https://keeprouter.com/models/doubao-embedding-vision-251215.md)
- [API observability](https://keeprouter.com/features/api-observability.md)
- [AI gateway vs direct provider APIs](https://keeprouter.com/compare/direct-provider-apis.md)

## Find a model that fits your task

Check model availability, input type and billing unit before choosing a service or writing integration code.

[Compare models and prices](https://keeprouter.com/models)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
