Built for retrieval teams
Build retrieval applications with an explicit embedding boundary
A RAG application needs a stable relationship between its document vectors, query vectors, retrieval index, and answer model. KeepRouter can provide catalog-listed embedding and generation routes, but the application owns its index, relevance tests, citations, and the decision to re-embed when the vector model changes. [1] [3] [4]
Last reviewed 2026-09-28 · Editorial review: KeepRouter Editorial
Start with the retrieval contract
Do not treat an embedding model as an interchangeable string in configuration. Document and query vectors must be produced with a compatible model, input preparation, dimensions, and distance metric. A new generation model can often be tested against the existing retrieved passages; a new embedding model usually requires a second vector space and re-embedding. Qdrant's migration guide describes parallel collections or named vectors for that change. KeepRouter supplies neither the vector database nor the relevance judgment.
A two-route architecture
| Stage | What to record | Owner |
|---|---|---|
| Ingest | Document ID, source revision, chunk rule, embedding model ID, vector dimensions | Your ingestion service |
| Query | Query embedding model ID, index version, filters, top-k and latency | Your retrieval service |
| Answer | Retrieved passage IDs, generation model ID, token usage and charge | Your application and gateway records |
| Review | Supported answer, unsupported answer, missing passage, permission failure | Your evaluation set |
Use the live model catalog for current KeepRouter model IDs and endpoints. The API contract separates text embeddings from a multimodal embedding route, but that does not mean both model types are currently listed. A model that accepts text, image, and video input may use /v1/embeddings/multimodal rather than /v1/embeddings. Inspect the exact model page and OpenAPI contract before sending production media. The multimodal embedding guide explains the input boundary; it does not claim that every input type shares one vector space with every other model.
Change an embedding model without a silent quality regression
Keep the old index serving while a new index is populated from the same source documents. Dual-write new or changed documents during backfill, compare corpus counts and deletion state, then run a fixed query set against both indexes. Record recall of required passages, permission-filter failures, answer support, p95 retrieval time and combined embedding plus generation cost. Switch a small cohort only after those gates pass. Keep the old index and model ID available for rollback until the new path is stable. The reindexing checklist gives the sequence and worksheet.
Control access and cost at the right layer
Keep the API key on the server. Use a model allowlist and spend limit for the embedding worker, and a separate key for generation. KeepRouter request records can show route, model, status, usage and charge; they cannot tell whether the right passage was retrieved. Join those records to your own document, query and outcome IDs without placing private document text in gateway metadata. Use the usage and observability guide for the record boundary.
If a required embedding model or vector capability is absent from the live catalog, use a direct provider route for that stage. A split architecture is better than claiming an unsupported model is portable through a gateway.
Frequently asked questions
Can I change the embedding model ID without reindexing?
Usually no. Build and evaluate a new vector space from the source documents; keep the old one for rollback.
Does KeepRouter host my vector database?
No. It exposes catalog-listed model routes; your application owns storage, indexing, access filters and retrieval evaluation.
Is the multimodal embedding route the same as /v1/embeddings?
No. Follow the exact endpoint on the selected model page. Some multimodal models use /v1/embeddings/multimodal.
What proves a RAG migration worked?
A fixed query set should pass passage recall, permission filters, answer support, latency and cost gates before cohort rollout.
Sources reviewed
Sources last reviewed 2026-09-28
- [1] Qdrant embedding migration tutorial
- [2] OpenAI embeddings guide
- [3] KeepRouter OpenAPI
- [4] KeepRouter live model catalog