Open WebUI Custom API: Connect Models Without Surprises

Add an OpenAI-compatible connection to Open WebUI, distinguish model discovery from generation, and keep chat, embeddings and extra tools correctly scoped.

Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Layered responsibility map between an application, AI gateway, and model providers
An application connects to a model API through an explicit request boundary. Conceptual illustration.

Add a custom model API in Open WebUI through the administrator's Connections settings, using the provider's base URL and key. Model discovery and chat generation are separate checks. A successful model list does not demonstrate tools, embeddings, image generation or every other feature offered by the interface.

This guide uses KeepRouter as a compatible chat destination. It is intended for an operator adding one connection to an existing Open WebUI instance, not for exposing a new unauthenticated service to the internet.

Add a small, explicit connection

The Open WebUI compatible-provider guide describes Settings, Admin and Connections. Create a separate connection rather than overwriting a working one. Use https://keeprouter.com/v1, a KeepRouter key and a small set of model IDs you intend to test.

Connection choiceTrial valueWhy
ProtocolOpenAI-compatible chatMatch the API contract
URLhttps://keeprouter.com/v1Let the client append operation paths
KeyStored administrator credentialKeep the secret out of shared prompts
Model IDsStart with free for textMake the first selection unambiguous
Extra capabilitiesAdd after validationAvoid testing several unknowns at once

Keep the default provider hint for a generic compatible service. Selecting a local-server-specific mode can expose management actions that a hosted API does not implement. A hosted endpoint cannot unload or download a local model merely because the interface offers such a control.

Verify discovery and generation independently

Open WebUI's documentation explains that connection verification uses the model-list endpoint. Some services need an explicit model allowlist when discovery is unavailable. Saving a connection and successfully generating an answer are different milestones.

Start a new chat, select the intended connection's free model, and ask a short question containing a checkable fact. Then inspect whether the response completes normally. If the model appears twice under different connections, rename the local display labels so you know which destination received the request.

Use a fixture such as “The project code is cedar-17; return that code exactly.” It tests the selected request path without depending on current world knowledge. The response does not establish coding, vision or reasoning quality. For a paid model, use its exact ID and a task-specific test after checking price and account balance.

Keep UI features tied to their real backend

Feature in the interfaceSeparate question
ChatDoes the selected route accept and stream messages?
File retrievalWhich service parses files and creates embeddings?
SpeechIs the required speech endpoint configured?
Image generationDoes the selected image API match this request format?
ToolsWho executes them and how are arguments authorized?
Conversation titlesWhich background model generates them?

Do not interpret a successful chat connection as support for all these features. A file upload can succeed while its retrieval pipeline fails later. A visible tool button can exist even when the model cannot produce the expected tool call.

Add one feature at a time and use a small fixture for each. For a document, ask about a unique sentence and inspect the retrieved text. For speech, use a short neutral phrase. For tools, start with a read-only operation rather than a command with external effects.

Know where requests originate

In the usual server-mediated setup, the Open WebUI backend reaches the API. A connection that works from your laptop can still fail inside a container with different DNS, proxy or certificate configuration. Inspect connectivity from the component that actually sends the request.

Open WebUI also documents Direct Connections, where browser requests go directly to a provider. That changes credential handling and browser-network requirements. Do not switch to that mode merely to avoid diagnosing a backend connection problem; decide whether that architecture is appropriate for your users.

For local servers, localhost means the current machine or container. That issue is separate from a public HTTPS API such as KeepRouter. Replacing a hosted API URL with a container hostname is not a general networking fix.

Account for background work

A chat interaction can trigger more work than the main visible response. Titles, tags, retrieval and other configured features may use separate models. Identify those settings before comparing a per-message estimate with your API usage.

For an illustrative budget, ten users sending twenty messages each create 200 visible messages. If each new conversation also generates a title, add those calls separately; do not assume one title per message. Use your actual conversation count and enabled features rather than multiplying by an arbitrary overhead percentage.

Compare useful completed conversations, not just successful HTTP responses. A model that answers quickly but ignores uploaded evidence can look healthy in a request chart while failing the user's task.

The live model catalog and quickstart establish supported IDs and basic access. Use the RAG cost guide when adding documents. Keep the original connection available until the new one passes both text generation and the specific extra features your team relies on.

Frequently asked questions

Does Verify Connection test generation?

It primarily checks the model-list connection described by Open WebUI. Send a separate chat request and validate the features your workflow uses.

Why does a browser test work but the app fail?

The server or container may use a different network path. Diagnose the component sending the API request rather than assuming the browser and backend share connectivity.

Can one chat connection power every media feature?

No. Speech, image generation and embeddings need suitable endpoints and models. Configure and test each feature independently.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Open WebUI compatible providers
  2. [2] Open WebUI direct connections

Related guides

← All posts · Models & pricing · Get an API key