LangChain Custom Base URL: An OpenAI-Compatible API Guide

Set ChatOpenAI base_url correctly, test streaming and tools separately, and avoid losing provider-specific fields when changing an LLM endpoint.

Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Field-by-field OpenAI-compatible API migration checklist from request contract to rollback
Validate the request contract and the application outcome when moving an integration. Conceptual illustration.

For a standard Chat Completions endpoint, create LangChain's ChatOpenAI with an explicit base_url, a key for that destination and its exact model ID. Keep provider-specific fields out of the initial request. A compatible transport does not guarantee that LangChain preserves every extended response field.

This guide is for a Python application that already uses LangChain and wants a controlled endpoint change. You do not need to rebuild the chain to discover whether the new service accepts the basic contract.

Make the client configuration explicit

The ChatOpenAI integration guide documents custom base URLs and warns that the wrapper targets standard OpenAI specifications. Non-standard reasoning fields may not survive normalization. Keep the raw API contract and the application-visible message format separate in your diagnosis.

Install langchain-openai in your application's isolated environment and record the resolved version in its lock file. Set KEEPROUTER_API_KEY outside the script. This example performs one real request against the free text model, with SDK retries disabled so a failed first attempt stays visible.

import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="free",
    base_url="https://keeprouter.com/v1",
    api_key=os.environ["KEEPROUTER_API_KEY"],
    use_responses_api=False,
    max_retries=0,
    timeout=30,
)
reply = llm.invoke([
    ("system", "Answer using only the provided fact."),
    ("human", "Fact: the release label is amber. What is the label?"),
])
print(reply.content)
print(reply.usage_metadata)

Expect an answer containing amber; do not require identical punctuation. Missing usage metadata is a separate result from missing text. Inspect the actual response and the wrapper version before filling a missing token count with zero. Zero is a measurement, while missing is an observation gap.

Distinguish routing from network proxying

base_url changes the API destination. It is not the same thing as an HTTP proxy used by your network. The destination still needs the right authentication and paths. Explicit configuration also avoids an unrelated OPENAI_BASE_URL or OPENAI_API_BASE environment variable changing the target between a local shell and production.

ValueControlsTypical mistake
base_urlThe model API origin and prefixAdding /chat/completions to a base URL
api_keyAuthentication at that originReusing an OpenAI key for another service
modelA destination-specific model IDCopying a router-specific namespace
HTTP proxy settingsHow your machine reaches the originTreating a proxy hostname as the model API

Keep the first request small and deterministic enough to inspect. Do not add a framework agent, retrieval pipeline and new model in the same change. If the request succeeds outside LangChain but fails inside it, compare serialized fields rather than changing the destination repeatedly.

Add one behavior at a time

After plain text, test llm.stream() with the same short prompt. Record whether the application receives incremental content and a normal finish. Add streaming usage only when the selected API path supports it. The integration documentation notes that usage behavior differs for custom endpoints; enabling an option in the client cannot force a provider to return a field.

Next test a harmless tool, such as a function that returns the opening hours of a fictional shop. Inspect the tool name, parsed arguments, ID and final answer after returning the tool result. A sentence that describes a tool call is not a tool call. The LangChain tool documentation is the framework reference; the selected model's route determines which capabilities are available.

Then test your real output parser. If a chain expects a string but the model returns structured content blocks, decide where normalization belongs. Avoid blindly converting every object to a string: that can make a test appear successful while downstream JSON validation fails.

Preserve information your application actually uses

If the application depends on provider-specific reasoning, citations, native search results or caching controls, write those requirements down before migration. Compare a redacted raw response with the AIMessage received by your chain. A missing field can disappear at the upstream API, gateway adaptation or framework wrapper.

Use a provider-specific integration when its extended contract is essential. A generic wrapper is useful precisely because it narrows the contract; it is a poor choice when the application needs fields outside that contract. This is a reason to keep a working direct integration for one feature while routing ordinary chat elsewhere.

Avoid retry multiplication in a chain

An illustrative chain with three model calls and two retries per call can make up to nine attempts. A task-level retry can increase that further. Choose which layer owns transient-error retries and keep a deadline for the whole user task. The failover guide covers the distinction between repeating a request and changing models.

For evaluation, log a non-sensitive case ID, selected model, finish state, latency, usage availability and whether the answer passed. Keep private prompts and credentials out of diagnostic output. Start with a dozen representative cases and inspect failures individually; that fixture size is useful for debugging, not a statistical quality claim.

Once the contract is stable, compare a paid model from the live catalog on the same cases. The migration checker helps inspect configuration without accepting your key. Keep the previous client configuration available until both the chain output and its cost per accepted result are understood.

Frequently asked questions

Should base_url end with chat/completions?

No. For this configuration it ends with /v1. The client appends the operation path; a full operation URL can create an invalid combined path.

Will ChatOpenAI preserve every reasoning field?

No. Its official scope is the OpenAI API contract. Check extended fields explicitly and use the relevant provider integration when those fields are required.

Why disable retries in the first test?

It exposes the first error and makes request counts easier to explain. Add bounded retries after diagnosing the request and deciding which layer owns them.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] LangChain ChatOpenAI
  2. [2] LangChain tools

Related guides

← All posts · Models & pricing · Get an API key