Dify Custom Model Setup with an OpenAI-Compatible API
Configure Dify model credentials, endpoint mode and capabilities; separate chat from embeddings and diagnose validation failures before publishing an app.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

Use Dify's OpenAI-API-compatible model provider to connect a compatible gateway, then configure the exact model, API base and credentials. Select capabilities supported by that model and route. Setting a checkbox for tools or structured output does not add those features to an API.
This walkthrough is for a team adding a model to an existing Dify app. The useful first milestone is a published test app answering a controlled question with the intended model, not simply a green credential dialog.
Separate the provider, model and app
Dify configures model access at workspace level. The model-provider guide explains the administrator role in that setup. An app then selects from configured models. If an administrator adds a new model but an existing workflow still points at the old one, changing the provider list has not migrated the workflow.
Install or select the official OpenAI-API-compatible provider available to your Dify version. Its published configuration schema distinguishes the model name, optional display label, API key and endpoint URL. Keep a record of the plugin version next to the workflow version so a future plugin update is not mistaken for a model change.
| Field | Initial KeepRouter value | Reason |
|---|---|---|
| Model name / endpoint model | free for a text connection test | Avoid borrowing another catalog's namespace |
| API base URL | https://keeprouter.com/v1 | The plugin constructs the operation URL |
| API key | A KeepRouter key stored in the provider form | Credentials must belong to the destination |
| Completion mode | Chat | Use Chat Completions for this example |
| Capabilities | Only those verified for the selected model | Advertised switches affect requests |
Use the actual fields exposed by your installed plugin. Display labels are for people; the endpoint model name controls the request. A friendly name such as “support assistant” should not accidentally replace a real model ID.
Build a three-node fixture
Create a small workflow with a text input, one LLM node and an output. Keep knowledge retrieval, tools and conversation memory out of the first experiment. Give the input a tiny policy document and ask a question with one checkable answer:
Policy: Trial accounts can create two projects. Paid accounts can create ten.
Question: How many projects can a trial account create?
Return the number and quote the sentence that supports it.The expected fact is two projects. The exact wording may vary. Reject an answer that invents a price, adds an account upgrade link, or answers for the paid tier. This tests whether the application receives the intended input and exposes a usable answer. It is not a benchmark for general reasoning.
Run the draft, then run the app's published version with the same input. If results differ, compare publication state and model selection before changing the prompt. A team can otherwise spend hours tuning a draft that users never receive.
Treat knowledge retrieval as a separate integration
A Dify knowledge base may need an embedding model and, depending on your design, a reranker. The chat model configured above does not automatically satisfy either role. The compatible provider supports several model types, but each still requires an appropriate endpoint and model.
Keep the existing embedding model while evaluating a new answer generator. That preserves the index and makes the comparison narrower. If you also change embeddings, review the reindexing guide; mixing vectors from different models in one index can make retrieval results meaningless.
For a retrieval fixture, add a second document with a similar sentence about a different account tier. Ask for both an answer and its source. Inspect the retrieved passages separately from the final response. Wrong context points to retrieval; correct context followed by a wrong answer points further downstream.
Find out what validation actually sent
Dify's compatible plugin performs credential-validation requests, and its implementation includes handling for endpoint and token-parameter differences. A validation failure can therefore be caused by the probe's request shape as well as authentication.
| Failure | Evidence to collect | Smallest useful fix |
|---|---|---|
| Unauthorized | Destination, key origin, status | Correct the provider credential |
| Unknown model | Actual endpoint model name | Copy a supported ID from the catalog |
| Invalid token parameter | Sanitized error and plugin version | Use the documented parameter mode |
| Unsupported response format | Requested schema feature | Disable only the unneeded feature or choose a supporting route |
| Works locally, fails in Dify | Connectivity from the plugin runtime | Fix that runtime's network path |
Do not turn off TLS verification to make a connection work. Avoid posting the credential or full private prompt when reporting a plugin issue. A minimal request shape, status and redacted error usually provide a better reproduction.
Evaluate the app's cost, not the connection label
An app can spend tokens on question rewriting, classification, retrieval summaries and its final answer. Count those model nodes and any reruns. An illustrative workflow with three model nodes and one repeated attempt makes four billable requests, not one. Measure the number of useful completed tasks and compare that to actual charges.
Before moving users, choose a suitable paid candidate from the model catalog, keep its limits in the Dify configuration, and run the same questions. Use the quickstart for account setup and the cost calculator for a first budget. Preserve the previous workflow version until the published app passes the fixture.
Frequently asked questions
Does changing a provider migrate every Dify app?
No. Check the model chosen in each workflow and its published version. Workspace availability and application selection are different settings.
Can the chat model also embed my knowledge base?
Only if an appropriate embedding model and endpoint are separately configured. A chat connection does not replace embedding or reranking configuration.
Is this a tested Dify performance benchmark?
No. This is a source-checked setup and validation guide. Run the supplied fixture in your installed Dify and plugin versions before drawing performance conclusions.
Sources reviewed
Article last reviewed 2026-09-29
- [1] Dify model providers
- [2] Official compatible-provider schema
- [3] Official compatible-provider implementation