Changelog
2026-10-03
More current models and clearer agent contracts
- Sixteen additional model IDs have prices and source-backed details: GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Claude Sonnet/Opus 5.5, GLM 5.3/Flash/FlashX, MiMo V2.6 Flash/Pro/Pro UltraSpeed, Grok 4.7, Doubao Seed 2.1 Pro/Turbo, Qwen 3.8 27B and Gemma 4 26B A4B IT. Doubao Pro selects the September 15 snapshot. Models and current prices
- Copyable examples use the appropriate API. GPT tool workflows start with Responses; Claude 5.5 starts with native Messages. Model JSON exposes reviewed endpoint, reasoning and pricing-policy guidance. Fixed customer prices for the relevant GPT, Claude and Grok IDs disclose their cache-write and long-context ceilings rather than presenting them as maker base rates.
- Five new guides in English and Chinese bring the library to 55 articles. They cover exact IDs, tool-history preservation, cost worksheets and task-based selection. MiMo's guide includes the announced October 21 V2.5 retirement. GPT guide · Claude migration · MiMo migration · GLM variants
- Kimi now lists K2.6, K2.7 Code, K2.7 Code High-Speed and K3 with separate input, output and cached-input prices. The model pages link to exact maker documentation and show verified context limits. K3's fixed input rate covers uncached input and one-hour cache writes; cache hits use the published cache rate.
- The quickstart and Gemini API guide clarify the Google Cloud chat route: use a KeepRouter key, a public Gemini ID and Chat Completions. Streaming is supported; Responses, Messages, Google-native generateContent and Live WebSockets are separate capabilities and are not exposed by this route.
What's new in KeepRouter. Have a request? Email support@keeprouter.com.
2026-09-29
More integration, troubleshooting and cost guides
- The library now contains 50 guides, each in English and Chinese. New walkthroughs cover n8n, Dify, LangChain, LlamaIndex, Cline, Continue, Aider, Open WebUI and LibreChat, plus migration from Roo Code. Each states its setup and capability limits. Browse the guides
- Reusable examples for difficult API decisions. Work through rate limits, interrupted streams, structured output, tool loops, context budgets, RAG and reasoning costs, document extraction, batch processing and local inference. Cost examples label their assumptions and link to current model prices. RAG cost worksheet · Document extraction evaluation
More Gemini models and clearer token usage
- Nine Gemini chat models are available. Seven model IDs join the catalog. All nine listed IDs passed bounded text calls through KeepRouter. The new guide includes exact IDs, current price links, an OpenAI SDK example and the tested streaming, JSON and tool-conversation scope. Gemini API guide
- Reasoning is counted once. Reported thinking tokens are included in completion usage, with the breakdown retained. Use each model page’s current input, output and cached-input rates to check cost. Models and prices
- More precise model attribution. HappyHorse pages now identify Alibaba, their task and resolution variants, and official examples. Text Embedding 004 identifies Google and separates its Cloud lifecycle from Gemini API retirement. HappyHorse I2V · Text Embedding 004
2026-09-13
DeepSeek V4 Pro continues after September 14
- The planned switchover is no longer happening. DeepSeek's current API notice says
deepseek-v4-prowill continue after September 14, 2026, and that any later change will receive a separate notice. KeepRouter therefore keeps V4 Pro as a separately callable, separately priced model; no production route or customer rate was changed for this notice. DeepSeek's notice · V4 Pro pricing & specs
Veo 3.1, completion webhooks, and per-model clip lengths
- Completion webhooks. Pass
webhook_urlwhen you submit a render and the gateway POSTs a signed event when the task finishes — no polling loop required. The event carries the same fields the poll returns, is signed withkr-signature(HMAC-SHA256 over<t>.<body>, using the secret returned once at submit), and is retried on a bounded schedule (about 1 min, 15 min, 1 h, 6 h) if your endpoint is down. Only public https destinations are accepted: internal, loopback and cloud-metadata addresses are refused before the task is ever submitted. The poll stays the source of truth and now reports the delivery state, so a broken endpoint is visible rather than silent. Video generation - Every poll response now carries
expires_at. The clip URL and the gateway's tracking window share one 24-hour lifetime; the field makes that machine-readable instead of leaving you to compute it. Video generation - 4K and 1080p tiers added across the video catalogue.
veo-3.1-fast-4k($0.30/s),veo-3.1-4k($0.60/s), and 1080p versions ofhappyhorse-1.0-i2v,happyhorse-1.0-r2vandhappyhorse-1.0-video-edit($0.24/s each). Check the selected model page for its current rate and supported resolution. Models & Pricing - Image-to-video, reference-to-video and video edit are published.
happyhorse-1.0-i2v,happyhorse-1.0-r2vandhappyhorse-1.0-video-editjoin the text-to-video tier at $0.14/s (720p). They take their input through amediaarray —{ "type": "first_frame" | "reference_image" | "video", "url": "https://…" }— where the URL must be publicly fetchable and a video input must be at least 3 seconds. Clip lengths are 3 to 15 seconds, refused locally when out of range. Video generation - The Veo 3.1 family is complete in both HD tiers.
veo-3.1-lite(720p $0.05/s, 1080p $0.08/s),veo-3.1-fast(720p $0.10/s, 1080p $0.12/s) andveo-3.1(720p and 1080p, $0.40/s). Completed test clips included synced audio. Models & Pricing - Veo 3.1 Lite is live on
POST /v1/video/generationsin two resolution tiers:veo-3.1-lite(720p) at $0.05/s andveo-3.1-lite-1080pat $0.08/s, each pinned so the request cannot be rendered above the tier it was priced at. At these customer rates, a 4-second clip costs $0.20 at 720p and $0.32 at 1080p. Models & Pricing - Clip lengths are checked per model before the render is submitted. Veo 3.1 accepts only 4, 6 or 8 seconds — it refuses anything else, and a min/max range would have let 5 and 7 through. The gateway now knows the difference between a range and a list, returns
400 invalid_durationwith the lengths the model does accept, and, when you omitduration, sends the shortest length that model supports (4 seconds for Veo, 5 for the rest) so "omitted" can never mean "rejected". - Cached-input rate corrected for Kimi K2.7 Code (High-Speed). Cached input is now $1.90 / 1M tokens, matching the uncached input rate. The published rates are $1.90 / 1M in, $8 / 1M out, $1.90 / 1M cached. Uncached input and output rates are unchanged. Models & Pricing
- Published prices receive regular checks. Pricing changes require review before publication. Check the exact model page for the current customer rate.
2026-09-12
Video generation, text-to-speech, and multimodal embeddings
- Video generation is live on
POST /v1/video/generations. Nineteen published model tiers across four families: Wan 2.7 (text-to-video, image-to-video, reference-to-video, and video edit, at 720p and 1080p), HappyHorse 1.0, PixVerse v3.5 through v6, and MiniMax H3. Video is an asynchronous task API — the submit call returns a task id and you collect the render fromGET /v1/videos/{id}(or/v1/video/generations/{id}; both work). Billing is per second of clip, charged when the task is accepted. How the flow works - Video prices are explicit for each resolution tier. Each priced resolution has its own model ID: for example,
wan2.7-t2vis 720p andwan2.7-t2v-1080pis 1080p. The gateway pins the selected tier so a request cannot be rendered at a higher resolution than the one priced. Models & Pricing is the current billing reference. - Tasks belong to the account that submitted them, and are tracked for 24 hours. A task's poll response carries a signed download URL, so it is readable only by the account that submitted it — another account gets the same
404as an unknown id. After 24 hours the gateway reports404 task_expiredrather than relaying a state nobody can act on. - Text-to-speech is published.
gemini-3.1-flash-tts-previewonPOST /v1/audio/speech, billed per second of generated audio at $0.0005/s. - Multimodal embeddings are published.
doubao-embedding-vision-250615and-251215embed text, images, and video onPOST /v1/embeddings/multimodal. Use the dedicated multimodal endpoint shown on the model pages. - Prices now carry a unit. A video model is priced per second of clip and a speech model per second of audio, so
/models, each model page, the Markdown twins, JSON-LD and/models.csvstate the unit instead of rendering a per-second rate as a per-million-token one./models.csvgainedoutput_price_unitandoutput_pricecolumns for the same reason. - New error codes for the video and audio routes:
resolution_not_priced,invalid_task_id,task_not_found,task_expired, andupstream_account_unavailable(returned when an upstream account cannot serve a model, instead of forwarding the provider's own account message). Each is documented at /docs/errors. - The catalog now lists 80 models. The npm SDK experiment was removed: a gateway should be reachable with the client libraries you already use, so the OpenAI- and Anthropic-compatible routes remain the whole integration surface.
2026-09-11
DeepSeek V4.1 Flash and four more current API models
- Five production-probed model IDs. KeepRouter now publishes
deepseek-flash,gpt-6-astra,claude-fable-5-1,gemini-3.8-flash, andqwen3.8-max-0902with maker-linked specifications and exact callable IDs. - Current fixed customer rates. Prices below are input / output / cached input per million tokens: DeepSeek Flash is $0.30 / $1.20 / $0.006 using the peak-hour rate as a fixed ceiling; GPT-6 Astra is $25 / $75 / $2, covering OpenAI's long-context cache-write ceiling; Claude Fable 5.1 is $20 / $50 / $0.25, covering the one-hour cache-write rate; Gemini 3.8 Flash is $0.75 / $3.75 / $0.075 through December 31, 2026; and Qwen3.8 Max 0902 is $2.50 / $6 / $0.25, covering cache writes. KeepRouter currently publishes one fixed input rate per model, so that same input rate applies to ordinary uncached input as well. Models & Pricing remains the billing source of truth.
- DeepSeek migration stays explicit. The retired
deepseek-v4-flashanddeepseek-v4-flash-vision-expIDs remain temporary compatibility aliases to V4.1 Flash and now use the same fixed rate. DeepSeek has since confirmed thatdeepseek-v4-prowill continue after September 14 with its billing method unchanged, so V4 Pro remains separately callable and priced. Current notice - Kimi K3 route repaired.
kimi-k3now uses a production-probed route without changing its public model ID or customer price. - Retired Kimi K2.5 removed. Moonshot retired
kimi-k2.5on August 31, so it is no longer published or routed as an available model. Requests should move to a current, explicitly tested Kimi model instead of relying on a silent alias. - Restricted variants stay out of the public catalog. Claude Mythos 5.1 and Gemini 3.8 Flash Cyber require vetted-access programs, so KeepRouter does not present them as generally callable models.
2026-08-26
DeepSeek Vision Exp and refreshed V4 rates
- New experimental multimodal model.
deepseek-v4-flash-vision-expaccepts mixed text and image input and returns text. DeepSeek documents a 1M-token context window, a 384K maximum output, and support for Chat Completions, Messages, and Responses. Its historical model page now resolves to the current DeepSeek V4.1 Flash page. - V4 prices refreshed. V4 Flash, V4 Pro and Vision Exp use the fixed customer rates shown in Models & Pricing. KeepRouter does not switch these rates by time of day.
- Callable ids stay explicit. Use
deepseek-v4-flash,deepseek-v4-pro, or the experimentaldeepseek-v4-flash-vision-exp. Backend version labels are not public model ids.
2026-08-15
A bilingual AI gateway decision library
- More useful comparisons. The comparison library now covers 15 buyer guides and product comparisons, including fit-based AI gateway and OpenRouter-alternative shortlists, managed versus self-hosted decisions, direct provider APIs, and seven additional gateway or cloud-platform comparisons. Shortlists are organized by operating fit rather than a synthetic ranking.
- Answers and production playbooks. Eight new direct-answer pages explain routing, failover, latency, security, BYOK, gateway selection, and adoption timing. Three new editorial guides provide security, latency-measurement, and failover test plans.
- Stable Simplified Chinese URLs. Every feature, audience, comparison, answer, and blog page now has a crawlable
/zh/...HTML and Markdown twin with localized schema, reciprocalhreflang, and language-preserving internal links. - Clearer reading and discovery. Long-form pages now use answer-first verdicts, sticky contents, decision tables, checklists, step cards, evidence panels, FAQ blocks, and complete related-guide grids. Sitemap,
llms-full.txt, AI index, RSS, SSR, and Markdown derive from the shared registries; rootllms.txtenumerates every English page and links the five Chinese content hubs.
2026-08-13
DeepSeek V4 Pro is generally available
deepseek-v4-pro now resolves to DeepSeek's GA release DeepSeek-V4-Pro-0813. The callable model id has not changed. Its catalog metadata now records the maker-documented 1M-token context window, 384K maximum output, and August 13 release date.
- The current stable DeepSeek ids are
deepseek-v4-proanddeepseek-v4-flash. The retired
deepseek-chat and deepseek-reasoner aliases are no longer suggested by the admin preset.
- DeepSeek now documents Chat Completions, Anthropic-compatible requests, and native Responses API
support for both V4 models. The admin includes separate Chat and Responses channel presets.
- Both V4 models default to thinking mode at high effort; the maker API supports low, high, and max.
2026-07-26
New: DeepSeek V4 Flash — a 1M-token context window
deepseek-v4-flash, a DeepSeek model with a maker-documented 1M-token context window, is now in the production catalog. Its historical link now resolves to DeepSeek V4.1 Flash pricing & specs.
- Also added:
deepseek-v4-pro, same maker-documented 1M-token context, on the chat route.
Current rates for both are published in Models & Pricing.
2026-07-17
- New: Kimi K3.
kimi-k3— Moonshot AI's flagship with a maker-documented 1M-token context window — is now on KeepRouter. See Kimi K3 pricing & specs. - More Kimi models. Added
kimi-k2.7-code-highspeed(coding-focused, high-speed serving tier) andkimi-k2.5, joining the existingkimi-k2.6. See Models & Pricing.
2026-07-12
- Public site redesigned. The homepage opens with a working request view backed by the production catalog and health probe. Models and prices use a product directory, and model pages pair facts with a copyable endpoint-specific request.
- Documentation and trust pages reorganized. Quickstart, errors, security, company, comparison, policies, changelog, glossary, and SDK guides use a shared three-column shell with persistent navigation and page contents.
- First-request path corrected. New examples default to
free. Public model CTAs preserve the selected model through email verification, and API key enable and disable actions send boolean state. - Purchase facts moved before signup. The public site now states the minimum card top-up, processing fee formula, tax handling, refund boundary, request-log fields, cache default, and exact health-probe scope.
2026-07-10 · v0.5.0
- Production operations complete. Production and staging are deployed separately; releases run through typecheck, Worker/frontend tests and build, then an isolated staging runtime/schema canary before production. Public staging keeps provider routing and self-serve OTP/email disabled by default rather than copying production credentials.
- Safer operator sessions. The browser admin now uses a signed HttpOnly, Secure, SameSite=Strict cookie; the master admin token remains a non-browser break-glass path and is no longer stored by frontend JavaScript.
- Hardened production sign-in. Transactional email delivery powers OTP, welcome and low-balance mail; a real bot challenge protects the production code-request flow. Staging's official test challenge is not treated as anti-bot protection.
- Production billing. One-time Paddle card top-ups, customer billing portal, signature-verified webhooks, idempotent crediting, and refund/chargeback clawbacks are live.
- Recovery path. Production D1 is exported on a schedule, with Worker rollback and D1 Time Travel documented for operators.
2026-07-05
- New: a free model.
freeis priced at $0 per token. It is intended for testing, prototyping, and evaluating the API before moving to a paid model. See Models & Pricing. - Mistral AI models. Added
mistral-large-latestandmistral-medium-latest; thecodestral-latestcode model and thedevstral-medium-latestcoding-agent model; themagistral-medium-latestandmagistral-small-latestreasoning models; theministral-3b-latest,ministral-8b-latestandministral-14b-latestedge models; and the open-weightopen-mistral-nemo. - Xiaomi MiMo & open-weight models. Added Xiaomi's
mimo-v2.5andmimo-v2.5-pro, plus the open-weightgpt-oss-120b(OpenAI) andgemma-4-31B-it(Google DeepMind). Current rates are published in Models & Pricing.
2026-07-04
- Claude Fable 5 is back.
claude-fable-5is available again on KeepRouter. See Models & Pricing.
2026-06-26
- Transparent top-up processing fee. Top-ups now show the selected credit amount and a separate card-processing fee before payment. Model usage is deducted at the published per-model rates, and the credited balance equals the amount selected.
- Activity log. The Usage page now has a filterable request log (by model & status) with a per-request detail view and CSV export.
- Billing history & invoices. A button on the Credits page opens the Paddle customer portal for invoices, receipts, and payment methods.
- Quickstart guide at /docs/quickstart and an API error reference.
- Low-balance email alerts send a one-time message after a paid balance crosses the warning threshold, plus a welcome email for new accounts.
2026-06-25
- New: Security & Data Handling page. The /security page documents request-log fields, upstream forwarding, and optional response caching. Prompt and completion bodies are not written to the API request log; response caching is off by default.
- Added Doubao models (ByteDance):
doubao-2.0-pro,doubao-2.0-mini,doubao-1.5-pro,doubao-1.5-lite. See Models & Pricing. - Published per-page metadata and structured data, added year-long immutable asset caching, and reduced the page bundle size.
2026-06-24
- Clearer image-model pricing. Image models now show a per-image price (e.g.
$0.04 / image) instead of a token rate. - Full bilingual site. English and 简体中文 are selected from the browser language and can be changed manually.
- New K-monogram logo and social preview cards.
2026-06-23
- Service status page at /status.
- Published Terms of Service, Privacy Policy, Acceptable Use Policy, and Refund Policy.
2026-06-22
- Credit & debit card top-ups. Paddle processes prepaid, pay-as-you-go top-ups as Merchant of Record. There is no monthly subscription.