DeepSeek V4.1 Flash API — pricing & specs
DeepSeek V4.1 Flash, made by DeepSeek, accepts text and image input and returns text. On KeepRouter, DeepSeek V4.1 Flash costs $0.3000 per 1M input tokens and $1.20 per 1M output tokens, billed pay-as-you-go with no monthly fee. Call it through a compatible KeepRouter endpoint supported by its active route, with the model id deepseek-flash.
| Maker | DeepSeek |
|---|---|
| Modality | Text, vision |
| Context window | 1,048,576 tokens |
| Max output | 393,216 tokens |
| Released | 2026-09-10 |
| Input price | $0.3000 per 1M tokens |
| Output price | $1.20 per 1M tokens |
| Cached input | $0.006000 per 1M tokens |
| Capabilities | Vision (image input), Chat completions, Streaming where supported, Tool calling where supported |
| Endpoint | POST /v1/chat/completions or POST /v1/messages (compatible chat routes) |
| Model id | deepseek-flash |
How pricing works for DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is billed per token — $0.3000 per 1M input tokens and $1.20 per 1M output tokens, with cached input at $0.006000 per 1M tokens. The published price is pay-as-you-go, with no monthly fee; actual request cost depends on measured token usage.
Calling DeepSeek V4.1 Flash on KeepRouter
Point a compatible client at the supported KeepRouter endpoint and set the model to deepseek-flash. KeepRouter preserves the client-facing request shape while handling upstream routing or translation. Use POST /v1/chat/completions or, on compatible chat routes, POST /v1/messages; streaming and tool support depend on the model's active route.
cURL
curl https://keeprouter.com/v1/chat/completions \
-H "Authorization: Bearer $KEEPROUTER_KEY" -H "Content-Type: application/json" \
-d '{"model":"deepseek-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this image?"},{"type":"image_url","image_url":{"url":"https://keeprouter.com/logo-512.png"}}]}]}'Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1", api_key=os.environ["KEEPROUTER_KEY"])
r = client.chat.completions.create(model="deepseek-flash", messages=[{"role":"user","content":[{"type":"text","text":"What is in this image?"},{"type":"image_url","image_url":{"url":"https://keeprouter.com/logo-512.png"}}]}])JavaScript
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://keeprouter.com/v1", apiKey: process.env.KEEPROUTER_KEY });
const r = await client.chat.completions.create({ model: "deepseek-flash", messages: [{ role: "user", content: [{ type: "text", text: "What is in this image?" }, { type: "image_url", image_url: { url: "https://keeprouter.com/logo-512.png" } }] }] });Estimate API costs
At the current KeepRouter customer price, an example workload of 1,000 total input tokens, no cached input, and 500 output tokens per request costs approximately $0.000900 per request. At 100 requests per day, that is $2.70 over 30 days. This is a usage estimate, excluding processing fees, taxes, retries and application infrastructure. Actual usage, cache hits and supported generation durations need their own checks.
Adjust quantities in the API cost calculator.
How to evaluate DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is listed on KeepRouter as text, vision under the exact id deepseek-flash. Use /v1/chat/completions for the listed route; a maker's upstream features do not automatically apply to this gateway endpoint. The published context window is 1,048,576 tokens and the published output limit is 393,216 tokens. These are limits, not a recommended request size.
First workload: Start with a text question plus one representative image, then repeat the same question with text only. Compare answer quality, latency and billed input/output tokens.
Before production: For agent use, test tool calls and structured output on this exact model route before production; an OpenAI-compatible chat endpoint alone does not prove either feature.
Agent setup paths
OpenCode
OpenCode documents a custom OpenAI-compatible provider. Set baseURL to https://keeprouter.com/v1 and list deepseek-flash as a model; validate tool behavior on the exact route. Official setup.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"keeprouter": {
"npm": "@ai-sdk/openai-compatible",
"name": "KeepRouter",
"options": {
"baseURL": "https://keeprouter.com/v1",
"apiKey": "{env:KEEPROUTER_API_KEY}"
},
"models": {
"deepseek-flash": {
"name": "DeepSeek V4.1 Flash"
}
}
}
}
}Continue
Continue documents provider: openai with a custom apiBase. Set apiBase to https://keeprouter.com/v1 and model to deepseek-flash; validate the selected feature and endpoint. Official setup.
name: KeepRouter
version: 0.0.1
schema: v1
models:
- name: DeepSeek V4.1 Flash
provider: openai
model: deepseek-flash
apiBase: https://keeprouter.com/v1
apiKey: <YOUR_KEEPROUTER_API_KEY>These are documented configuration paths; model-specific tool, streaming and multimodal behavior still needs a real request test.
Public model usage evidence
OpenRouter identifies the corresponding variant as deepseek/deepseek-v4.1-flash. The figures below are OpenRouter token volume, not API call counts, KeepRouter traffic or market-wide usage.
- Model total: 18.6T prompt + completion tokens over 7 days through 2026-09-26. OpenRouter ranking.
- Hermes Agent (agent): 5.2T tokens attributed to this exact model in OpenRouter's public Apps block; observed 2026-09-28. The model page does not state the Apps block's measurement window. OpenRouter model apps.
- pi (agent): 2.07T tokens attributed to this exact model in OpenRouter's public Apps block; observed 2026-09-28. The model page does not state the Apps block's measurement window. OpenRouter model apps.
App attribution is opt-in and reflects traffic through OpenRouter. It does not prove the app uses KeepRouter, endorse this gateway, or establish a model quality ranking. Tokenizers and windows can differ; do not add the weekly total to the app figures.
Source: OpenRouter rankings through 2026-09-26 and model Apps observed 2026-09-28. Rankings data CC BY 4.0.
Published examples and cases
OpenRouter's model-specific Apps view publicly attributes traffic to Hermes Agent and pi. This is a public adoption signal, not a published implementation study or a KeepRouter customer claim.
Source and verification boundary
Official DeepSeek V4.1 Flash documentation. KeepRouter's live catalog is authoritative for the customer price and enabled endpoint shown here; the maker remains authoritative for upstream model capabilities and limits.
Pricing and implementation guides
- DeepSeek API pricing and cached-token cost
- Calculate your API workload cost
- KeepRouter API keys: move from free to paid models
Guides
Related models
- DeepSeek V3.1 by DeepSeek — $0.2100 per 1M input tokens and $0.7900 per 1M output tokens
- DeepSeek V4 Pro by DeepSeek — $1.32 per 1M input tokens and $3.96 per 1M output tokens
- DeepSeek V3.2 by DeepSeek — $0.2288 per 1M input tokens and $0.3432 per 1M output tokens
- Codestral by Mistral AI — $0.3000 per 1M input tokens and $0.9000 per 1M output tokens
- Claude Sonnet 5.5 by Anthropic — $4.00 per 1M input tokens and $10.00 per 1M output tokens
- Claude Sonnet 4.6 by Anthropic — $3.00 per 1M input tokens and $15.00 per 1M output tokens
All models & pricing · Quickstart · KeepRouter vs OpenRouter · Glossary · Get an API key
Catalog facts and prices last changed .
Page content reviewed .