# Build retrieval applications with an explicit embedding boundary

> A RAG application needs a stable relationship between its document vectors, query vectors, retrieval index, and answer model. KeepRouter can provide catalog-listed embedding and generation routes, but the application owns its index, relevance tests, citations, and the decision to re-embed when the vector model changes. [1](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/) [3](https://keeprouter.com/api/openapi.json) [4](https://keeprouter.com/models)

_Last reviewed 2026-09-28 · [Editorial review](https://keeprouter.com/editorial-policy#editorial-team)_

## Start with the retrieval contract

Do not treat an embedding model as an interchangeable string in configuration. Document and query vectors must be produced with a compatible model, input preparation, dimensions, and distance metric. A new generation model can often be tested against the existing retrieved passages; a new embedding model usually requires a second vector space and re-embedding. Qdrant's [migration guide](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/) describes parallel collections or named vectors for that change. KeepRouter supplies neither the vector database nor the relevance judgment.

## A two-route architecture

| Stage | What to record | Owner |
| --- | --- | --- |
| Ingest | Document ID, source revision, chunk rule, embedding model ID, vector dimensions | Your ingestion service |
| Query | Query embedding model ID, index version, filters, top-k and latency | Your retrieval service |
| Answer | Retrieved passage IDs, generation model ID, token usage and charge | Your application and gateway records |
| Review | Supported answer, unsupported answer, missing passage, permission failure | Your evaluation set |

Use the [live model catalog](/models) for current KeepRouter model IDs and endpoints. The API contract separates text embeddings from a multimodal embedding route, but that does not mean both model types are currently listed. A model that accepts text, image, and video input may use /v1/embeddings/multimodal rather than /v1/embeddings. Inspect the exact [model page](/models/doubao-embedding-vision-251215) and [OpenAPI contract](https://keeprouter.com/api/openapi.json) before sending production media. The [multimodal embedding guide](/blog/multimodal-embeddings-guide) explains the input boundary; it does not claim that every input type shares one vector space with every other model.

## Change an embedding model without a silent quality regression

Keep the old index serving while a new index is populated from the same source documents. Dual-write new or changed documents during backfill, compare corpus counts and deletion state, then run a fixed query set against both indexes. Record recall of required passages, permission-filter failures, answer support, p95 retrieval time and combined embedding plus generation cost. Switch a small cohort only after those gates pass. Keep the old index and model ID available for rollback until the new path is stable. The [reindexing checklist](/blog/embedding-model-migration-reindex-checklist) gives the sequence and worksheet.

## Control access and cost at the right layer

Keep the API key on the server. Use a model allowlist and spend limit for the embedding worker, and a separate key for generation. KeepRouter request records can show route, model, status, usage and charge; they cannot tell whether the right passage was retrieved. Join those records to your own document, query and outcome IDs without placing private document text in gateway metadata. Use the [usage and observability guide](/features/api-observability) for the record boundary.

If a required embedding model or vector capability is absent from the live catalog, use a direct provider route for that stage. A split architecture is better than claiming an unsupported model is portable through a gateway.

## Frequently asked questions

### Can I change the embedding model ID without reindexing?

Usually no. Build and evaluate a new vector space from the source documents; keep the old one for rollback.

### Does KeepRouter host my vector database?

No. It exposes catalog-listed model routes; your application owns storage, indexing, access filters and retrieval evaluation.

### Is the multimodal embedding route the same as /v1/embeddings?

No. Follow the exact endpoint on the selected model page. Some multimodal models use /v1/embeddings/multimodal.

### What proves a RAG migration worked?

A fixed query set should pass passage recall, permission filters, answer support, latency and cost gates before cohort rollout.

## Sources reviewed

_Sources last reviewed 2026-09-28_

1. [Qdrant embedding migration tutorial](https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/)
2. [OpenAI embeddings guide](https://developers.openai.com/api/docs/guides/embeddings)
3. [KeepRouter OpenAPI](https://keeprouter.com/api/openapi.json)
4. [KeepRouter live model catalog](https://keeprouter.com/models)

## Related guides

- [Embedding model migration: reindex without losing retrieval quality](https://keeprouter.com/blog/embedding-model-migration-reindex-checklist.md)
- [Multimodal embeddings: one vector space for text, images, and video](https://keeprouter.com/blog/multimodal-embeddings-guide.md)
- [doubao embedding vision 251215](https://keeprouter.com/models/doubao-embedding-vision-251215.md)
- [API observability](https://keeprouter.com/features/api-observability.md)
- [AI gateway vs direct provider APIs](https://keeprouter.com/compare/direct-provider-apis.md)

## Test one retrieval path

Choose an exact published embedding route, index a small permitted corpus, and measure retrieval before connecting answer generation.

[Create a free key](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree) · [Live models and pricing](https://keeprouter.com/models.md)
