# LlamaIndex OpenAILike: Change the LLM, Keep the RAG Index

> Configure LlamaIndex OpenAILike without silently changing embeddings. Test generation and retrieval separately with a small, reproducible RAG fixture.

_Published 2026-09-29 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 4 minute read_

![Three typed inputs (text, image, video) converging into one vector block](https://keeprouter.com/editorial/blog/multimodal-embeddings-guide.png)

_Keep retrieval evidence and answer generation observable as separate stages. Conceptual illustration._

Use LlamaIndex's `OpenAILike` integration for a compatible chat endpoint, and keep the embedding configuration explicit. Changing the answer-generating model does not require rebuilding an otherwise compatible vector index. Changing the embedding model often does. Treat those as separate experiments.

This guide is for an existing retrieval application whose team wants to try another LLM without losing its retrieval baseline. The first question is whether the new generator can answer from the same passages, not whether a completely rebuilt pipeline produces a different answer.

## Establish the two model boundaries

A typical RAG request embeds a question, retrieves passages and sends those passages to a generator. The document vectors and query vectors must occupy the same embedding space. The generator reads text, so its identity is a separate choice from the model that created those vectors.

| Component | Keep fixed for a generator test | Evidence to save |
| --- | --- | --- |
| Document processing | Same parsing and chunking | Document revision and chunk IDs |
| Embeddings | Same model and dimensions | Index configuration |
| Retrieval | Same query and ranking parameters | Retrieved passage IDs |
| Generation | Change only this model and endpoint | Exact prompt, answer and usage |

Avoid an all-at-once switch in global settings. A default model left implicit can start calling another service when a library configuration changes. If your application uses shared settings, inspect both `Settings.llm` and `Settings.embed_model`, following the [LlamaIndex settings reference](https://developers.llamaindex.ai/python/framework/module_guides/supporting_modules/settings/).

## Configure OpenAILike for the endpoint

The [OpenAILike API reference](https://developers.llamaindex.ai/python/framework-api-reference/llms/openai_like/) documents `api_base`, `is_chat_model` and the separate function-calling capability flag. Install `llama-index-llms-openai-like` in the application environment. This standalone example sends one text request; it does not construct an index or invoke an embedding model.

```python
import os
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="free",
    api_base="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_API_KEY"],
    is_chat_model=True,
    is_function_calling_model=False,
    context_window=4096,
    max_tokens=256,
    max_retries=0,
)
result = llm.complete(
    "Source A: Warehouse Cedar closes at 18:00. "
    "Question: When does Cedar close? Answer with the source label."
)
print(result.text)
```

The 4,096-token context and 256-token output are conservative application settings for this small fixture, not published limits of the free route. For a paid model, replace them with values supported by that exact model and keep room for the system prompt. Setting a larger context number in the client does not enlarge the service's actual window.

## Reuse the index and compare the retrieved passages

In your existing application, pass the new LLM explicitly to the query engine where supported, while leaving the stored index and embedding configuration unchanged. A common pattern is `index.as_query_engine(llm=llm)`. Load the existing index through the same storage and embedding setup already used in production.

Create a tiny document fixture with Cedar closing at 18:00 and Birch closing at 20:00. Include one question about a warehouse that does not appear. The expected behavior is a correct answer for Cedar, a different correct answer for Birch, and an admission that the third fact is missing. These are authored test cases, not a model benchmark.

Record `source_nodes` or the equivalent retrieved evidence exposed by your query engine. If Birch's passage is supplied for a Cedar question, repair retrieval before grading the generator. If the correct passage is supplied but the answer gives 20:00, examine prompt assembly and generation. This separation prevents a model upgrade from disguising an indexing problem.

## Define what counts as a supported citation

A bracketed source label alone is not enough. Check that the label identifies a retrieved passage and that the quoted text actually supports the claim. A generator can cite a real passage for an unsupported conclusion, especially when several documents have similar headings.

For each fixture, store the expected fact and acceptable supporting passage IDs outside the model prompt. Let a reviewer inspect disagreement. Do not ask the same model to generate an answer and then treat its own confidence score as proof. The [retrieval migration guide](/blog/embedding-model-migration-reindex-checklist) explains when a separate index is needed for an embedding change.

## Budget the whole RAG query

Generation input includes instructions, the question, conversation history and retrieved text. Suppose a hypothetical query includes 600 tokens of instructions, 100 of question text, 2,400 of passages and 900 of history. That is 4,000 input tokens before the answer. Reducing duplicate passages may matter more than changing the query-embedding price.

Keep ingestion and query costs separate. Re-embedding a library is an occasional cost; repeatedly sending unnecessary passages is a recurring one. Count query rewriting and reranking if your pipeline uses them. See the [RAG cost worksheet](/blog/rag-api-cost-per-query) for an example that includes retries and rejected answers.

## Move one workload before the whole library

Run the same questions through the old and new generators, with the same retrieval results. Compare supported answers, abstentions, latency and total charges. Preserve the previous LLM configuration so rollback does not require rebuilding the index.

Check the [live model catalog](/models) for a suitable paid route and use the [quickstart](/docs/quickstart) to establish basic access. A free connectivity result is useful for configuration; it does not establish that a model can handle your full document collection, long contexts or native tools.

## Frequently asked questions

### Does changing the LLM require re-embedding documents?

Not when only the text generator changes and the existing embedding model, dimensions and retrieval contract stay the same. An embedding change needs its own index plan.

### Why set is_chat_model explicitly?

It tells the wrapper to use chat behavior. It does not detect or grant the endpoint's capabilities; function calling has its own setting and must be verified separately.

### Is a correct-looking citation enough?

No. Verify that the cited passage exists and supports the specific answer. Matching a source label without checking its content can hide retrieval or reasoning errors.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [LlamaIndex OpenAILike reference](https://developers.llamaindex.ai/python/framework-api-reference/llms/openai_like/)
2. [LlamaIndex Settings](https://developers.llamaindex.ai/python/framework/module_guides/supporting_modules/settings/)

## Related guides

- [LangChain Custom Base URL: An OpenAI-Compatible API Guide](https://keeprouter.com/blog/langchain-openai-compatible-api.md)
- [RAG Cost per Query: A Worksheet Beyond Token Prices](https://keeprouter.com/blog/rag-api-cost-per-query.md)
- [Embedding model migration: reindex without losing retrieval quality](https://keeprouter.com/blog/embedding-model-migration-reindex-checklist.md)

## Try the setup with one small request

Follow the configuration steps, keep the key on your server, and check the returned answer and usage.

[Open the setup guide](https://keeprouter.com/docs/quickstart)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
