Dify Custom Model Setup with an OpenAI-Compatible API

Configure Dify model credentials, endpoint mode and capabilities; separate chat from embeddings and diagnose validation failures before publishing an app.

Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Layered responsibility map between an application, AI gateway, and model providers
An application connects to a model API through an explicit request boundary. Conceptual illustration.

Use Dify's OpenAI-API-compatible model provider to connect a compatible gateway, then configure the exact model, API base and credentials. Select capabilities supported by that model and route. Setting a checkbox for tools or structured output does not add those features to an API.

This walkthrough is for a team adding a model to an existing Dify app. The useful first milestone is a published test app answering a controlled question with the intended model, not simply a green credential dialog.

Separate the provider, model and app

Dify configures model access at workspace level. The model-provider guide explains the administrator role in that setup. An app then selects from configured models. If an administrator adds a new model but an existing workflow still points at the old one, changing the provider list has not migrated the workflow.

Install or select the official OpenAI-API-compatible provider available to your Dify version. Its published configuration schema distinguishes the model name, optional display label, API key and endpoint URL. Keep a record of the plugin version next to the workflow version so a future plugin update is not mistaken for a model change.

FieldInitial KeepRouter valueReason
Model name / endpoint modelfree for a text connection testAvoid borrowing another catalog's namespace
API base URLhttps://keeprouter.com/v1The plugin constructs the operation URL
API keyA KeepRouter key stored in the provider formCredentials must belong to the destination
Completion modeChatUse Chat Completions for this example
CapabilitiesOnly those verified for the selected modelAdvertised switches affect requests

Use the actual fields exposed by your installed plugin. Display labels are for people; the endpoint model name controls the request. A friendly name such as “support assistant” should not accidentally replace a real model ID.

Build a three-node fixture

Create a small workflow with a text input, one LLM node and an output. Keep knowledge retrieval, tools and conversation memory out of the first experiment. Give the input a tiny policy document and ask a question with one checkable answer:

Policy: Trial accounts can create two projects. Paid accounts can create ten.
Question: How many projects can a trial account create?
Return the number and quote the sentence that supports it.

The expected fact is two projects. The exact wording may vary. Reject an answer that invents a price, adds an account upgrade link, or answers for the paid tier. This tests whether the application receives the intended input and exposes a usable answer. It is not a benchmark for general reasoning.

Run the draft, then run the app's published version with the same input. If results differ, compare publication state and model selection before changing the prompt. A team can otherwise spend hours tuning a draft that users never receive.

Treat knowledge retrieval as a separate integration

A Dify knowledge base may need an embedding model and, depending on your design, a reranker. The chat model configured above does not automatically satisfy either role. The compatible provider supports several model types, but each still requires an appropriate endpoint and model.

Keep the existing embedding model while evaluating a new answer generator. That preserves the index and makes the comparison narrower. If you also change embeddings, review the reindexing guide; mixing vectors from different models in one index can make retrieval results meaningless.

For a retrieval fixture, add a second document with a similar sentence about a different account tier. Ask for both an answer and its source. Inspect the retrieved passages separately from the final response. Wrong context points to retrieval; correct context followed by a wrong answer points further downstream.

Find out what validation actually sent

Dify's compatible plugin performs credential-validation requests, and its implementation includes handling for endpoint and token-parameter differences. A validation failure can therefore be caused by the probe's request shape as well as authentication.

FailureEvidence to collectSmallest useful fix
UnauthorizedDestination, key origin, statusCorrect the provider credential
Unknown modelActual endpoint model nameCopy a supported ID from the catalog
Invalid token parameterSanitized error and plugin versionUse the documented parameter mode
Unsupported response formatRequested schema featureDisable only the unneeded feature or choose a supporting route
Works locally, fails in DifyConnectivity from the plugin runtimeFix that runtime's network path

Do not turn off TLS verification to make a connection work. Avoid posting the credential or full private prompt when reporting a plugin issue. A minimal request shape, status and redacted error usually provide a better reproduction.

Evaluate the app's cost, not the connection label

An app can spend tokens on question rewriting, classification, retrieval summaries and its final answer. Count those model nodes and any reruns. An illustrative workflow with three model nodes and one repeated attempt makes four billable requests, not one. Measure the number of useful completed tasks and compare that to actual charges.

Before moving users, choose a suitable paid candidate from the model catalog, keep its limits in the Dify configuration, and run the same questions. Use the quickstart for account setup and the cost calculator for a first budget. Preserve the previous workflow version until the published app passes the fixture.

Frequently asked questions

Does changing a provider migrate every Dify app?

No. Check the model chosen in each workflow and its published version. Workspace availability and application selection are different settings.

Can the chat model also embed my knowledge base?

Only if an appropriate embedding model and endpoint are separately configured. A chat connection does not replace embedding or reranking configuration.

Is this a tested Dify performance benchmark?

No. This is a source-checked setup and validation guide. Run the supplied fixture in your installed Dify and plugin versions before drawing performance conclusions.

Sources reviewed

Article last reviewed 2026-09-29

  1. [1] Dify model providers
  2. [2] Official compatible-provider schema
  3. [3] Official compatible-provider implementation

Related guides

← All posts · Models & pricing · Get an API key