# LangChain Custom Base URL: An OpenAI-Compatible API Guide

> Set ChatOpenAI base_url correctly, test streaming and tools separately, and avoid losing provider-specific fields when changing an LLM endpoint.

_Published 2026-09-29 · Updated 2026-09-29 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 5 minute read_

![Field-by-field OpenAI-compatible API migration checklist from request contract to rollback](https://keeprouter.com/editorial/blog/openai-compatible-api-migration-checklist.png)

_Validate the request contract and the application outcome when moving an integration. Conceptual illustration._

For a standard Chat Completions endpoint, create LangChain's `ChatOpenAI` with an explicit `base_url`, a key for that destination and its exact model ID. Keep provider-specific fields out of the initial request. A compatible transport does not guarantee that LangChain preserves every extended response field.

This guide is for a Python application that already uses LangChain and wants a controlled endpoint change. You do not need to rebuild the chain to discover whether the new service accepts the basic contract.

## Make the client configuration explicit

The [ChatOpenAI integration guide](https://docs.langchain.com/oss/python/integrations/chat/openai) documents custom base URLs and warns that the wrapper targets standard OpenAI specifications. Non-standard reasoning fields may not survive normalization. Keep the raw API contract and the application-visible message format separate in your diagnosis.

Install `langchain-openai` in your application's isolated environment and record the resolved version in its lock file. Set `KEEPROUTER_API_KEY` outside the script. This example performs one real request against the free text model, with SDK retries disabled so a failed first attempt stays visible.

```python
import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="free",
    base_url="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_API_KEY"],
    use_responses_api=False,
    max_retries=0,
    timeout=30,
)
reply = llm.invoke([
    ("system", "Answer using only the provided fact."),
    ("human", "Fact: the release label is amber. What is the label?"),
])
print(reply.content)
print(reply.usage_metadata)
```

Expect an answer containing amber; do not require identical punctuation. Missing usage metadata is a separate result from missing text. Inspect the actual response and the wrapper version before filling a missing token count with zero. Zero is a measurement, while missing is an observation gap.

## Distinguish routing from network proxying

`base_url` changes the API destination. It is not the same thing as an HTTP proxy used by your network. The destination still needs the right authentication and paths. Explicit configuration also avoids an unrelated `OPENAI_BASE_URL` or `OPENAI_API_BASE` environment variable changing the target between a local shell and production.

| Value | Controls | Typical mistake |
| --- | --- | --- |
| `base_url` | The model API origin and prefix | Adding `/chat/completions` to a base URL |
| `api_key` | Authentication at that origin | Reusing an OpenAI key for another service |
| `model` | A destination-specific model ID | Copying a router-specific namespace |
| HTTP proxy settings | How your machine reaches the origin | Treating a proxy hostname as the model API |

Keep the first request small and deterministic enough to inspect. Do not add a framework agent, retrieval pipeline and new model in the same change. If the request succeeds outside LangChain but fails inside it, compare serialized fields rather than changing the destination repeatedly.

## Add one behavior at a time

After plain text, test `llm.stream()` with the same short prompt. Record whether the application receives incremental content and a normal finish. Add streaming usage only when the selected API path supports it. The integration documentation notes that usage behavior differs for custom endpoints; enabling an option in the client cannot force a provider to return a field.

Next test a harmless tool, such as a function that returns the opening hours of a fictional shop. Inspect the tool name, parsed arguments, ID and final answer after returning the tool result. A sentence that describes a tool call is not a tool call. The [LangChain tool documentation](https://docs.langchain.com/oss/python/langchain/tools) is the framework reference; the selected model's route determines which capabilities are available.

Then test your real output parser. If a chain expects a string but the model returns structured content blocks, decide where normalization belongs. Avoid blindly converting every object to a string: that can make a test appear successful while downstream JSON validation fails.

## Preserve information your application actually uses

If the application depends on provider-specific reasoning, citations, native search results or caching controls, write those requirements down before migration. Compare a redacted raw response with the `AIMessage` received by your chain. A missing field can disappear at the upstream API, gateway adaptation or framework wrapper.

Use a provider-specific integration when its extended contract is essential. A generic wrapper is useful precisely because it narrows the contract; it is a poor choice when the application needs fields outside that contract. This is a reason to keep a working direct integration for one feature while routing ordinary chat elsewhere.

## Avoid retry multiplication in a chain

An illustrative chain with three model calls and two retries per call can make up to nine attempts. A task-level retry can increase that further. Choose which layer owns transient-error retries and keep a deadline for the whole user task. The [failover guide](/blog/llm-failover-design-guide) covers the distinction between repeating a request and changing models.

For evaluation, log a non-sensitive case ID, selected model, finish state, latency, usage availability and whether the answer passed. Keep private prompts and credentials out of diagnostic output. Start with a dozen representative cases and inspect failures individually; that fixture size is useful for debugging, not a statistical quality claim.

Once the contract is stable, compare a paid model from the [live catalog](/models) on the same cases. The [migration checker](/tools/api-migration-checker) helps inspect configuration without accepting your key. Keep the previous client configuration available until both the chain output and its cost per accepted result are understood.

## Frequently asked questions

### Should base_url end with chat/completions?

No. For this configuration it ends with /v1. The client appends the operation path; a full operation URL can create an invalid combined path.

### Will ChatOpenAI preserve every reasoning field?

No. Its official scope is the OpenAI API contract. Check extended fields explicitly and use the relevant provider integration when those fields are required.

### Why disable retries in the first test?

It exposes the first error and makes request counts easier to explain. Add bounded retries after diagnosing the request and deciding which layer owns them.

## Sources reviewed

_Article last reviewed 2026-09-29_

1. [LangChain ChatOpenAI](https://docs.langchain.com/oss/python/integrations/chat/openai)
2. [LangChain tools](https://docs.langchain.com/oss/python/langchain/tools)

## Related guides

- [LlamaIndex OpenAILike: Change the LLM, Keep the RAG Index](https://keeprouter.com/blog/llamaindex-openai-like-rag.md)
- [LLM Tool Calling: Build the Full Request-Result Loop](https://keeprouter.com/blog/llm-tool-calling-loop.md)
- [api migration checker](https://keeprouter.com/tools/api-migration-checker)

## Check your request before migrating

Check model IDs and request fields in your browser. No API key or inference call is needed.

[Check a sample configuration](https://keeprouter.com/tools/api-migration-checker)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
