Open WebUI Custom API: Connect Models Without Surprises
Add an OpenAI-compatible connection to Open WebUI, distinguish model discovery from generation, and keep chat, embeddings and extra tools correctly scoped.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Add a custom model API in Open WebUI through the administrator's Connections settings, using the provider's base URL and key. Model discovery and chat generation are separate checks. A successful model list does not demonstrate tools, embeddings, image generation or every other feature offered by the interface.
This guide uses KeepRouter as a compatible chat destination. It is intended for an operator adding one connection to an existing Open WebUI instance, not for exposing a new unauthenticated service to the internet.
Add a small, explicit connection
The Open WebUI compatible-provider guide describes Settings, Admin and Connections. Create a separate connection rather than overwriting a working one. Use https://keeprouter.com/v1, a KeepRouter key and a small set of model IDs you intend to test.
| Connection choice | Trial value | Why |
|---|---|---|
| Protocol | OpenAI-compatible chat | Match the API contract |
| URL | https://keeprouter.com/v1 | Let the client append operation paths |
| Key | Stored administrator credential | Keep the secret out of shared prompts |
| Model IDs | Start with free for text | Make the first selection unambiguous |
| Extra capabilities | Add after validation | Avoid testing several unknowns at once |
Keep the default provider hint for a generic compatible service. Selecting a local-server-specific mode can expose management actions that a hosted API does not implement. A hosted endpoint cannot unload or download a local model merely because the interface offers such a control.
Verify discovery and generation independently
Open WebUI's documentation explains that connection verification uses the model-list endpoint. Some services need an explicit model allowlist when discovery is unavailable. Saving a connection and successfully generating an answer are different milestones.
Start a new chat, select the intended connection's free model, and ask a short question containing a checkable fact. Then inspect whether the response completes normally. If the model appears twice under different connections, rename the local display labels so you know which destination received the request.
Use a fixture such as “The project code is cedar-17; return that code exactly.” It tests the selected request path without depending on current world knowledge. The response does not establish coding, vision or reasoning quality. For a paid model, use its exact ID and a task-specific test after checking price and account balance.
Keep UI features tied to their real backend
| Feature in the interface | Separate question |
|---|---|
| Chat | Does the selected route accept and stream messages? |
| File retrieval | Which service parses files and creates embeddings? |
| Speech | Is the required speech endpoint configured? |
| Image generation | Does the selected image API match this request format? |
| Tools | Who executes them and how are arguments authorized? |
| Conversation titles | Which background model generates them? |
Do not interpret a successful chat connection as support for all these features. A file upload can succeed while its retrieval pipeline fails later. A visible tool button can exist even when the model cannot produce the expected tool call.
Add one feature at a time and use a small fixture for each. For a document, ask about a unique sentence and inspect the retrieved text. For speech, use a short neutral phrase. For tools, start with a read-only operation rather than a command with external effects.
Know where requests originate
In the usual server-mediated setup, the Open WebUI backend reaches the API. A connection that works from your laptop can still fail inside a container with different DNS, proxy or certificate configuration. Inspect connectivity from the component that actually sends the request.
Open WebUI also documents Direct Connections, where browser requests go directly to a provider. That changes credential handling and browser-network requirements. Do not switch to that mode merely to avoid diagnosing a backend connection problem; decide whether that architecture is appropriate for your users.
For local servers, localhost means the current machine or container. That issue is separate from a public HTTPS API such as KeepRouter. Replacing a hosted API URL with a container hostname is not a general networking fix.
Account for background work
A chat interaction can trigger more work than the main visible response. Titles, tags, retrieval and other configured features may use separate models. Identify those settings before comparing a per-message estimate with your API usage.
For an illustrative budget, ten users sending twenty messages each create 200 visible messages. If each new conversation also generates a title, add those calls separately; do not assume one title per message. Use your actual conversation count and enabled features rather than multiplying by an arbitrary overhead percentage.
Compare useful completed conversations, not just successful HTTP responses. A model that answers quickly but ignores uploaded evidence can look healthy in a request chart while failing the user's task.
The live model catalog and quickstart establish supported IDs and basic access. Use the RAG cost guide when adding documents. Keep the original connection available until the new one passes both text generation and the specific extra features your team relies on.
Frequently asked questions
Does Verify Connection test generation?
It primarily checks the model-list connection described by Open WebUI. Send a separate chat request and validate the features your workflow uses.
Why does a browser test work but the app fail?
The server or container may use a different network path. Diagnose the component sending the API request rather than assuming the browser and backend share connectivity.
Can one chat connection power every media feature?
No. Speech, image generation and embeddings need suitable endpoints and models. Configure and test each feature independently.
Sources reviewed
Article last reviewed 2026-09-29