Changelog

2026-10-03

More current models and clearer agent contracts

  • Sixteen additional model IDs have prices and source-backed details: GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, Claude Sonnet/Opus 5.5, GLM 5.3/Flash/FlashX, MiMo V2.6 Flash/Pro/Pro UltraSpeed, Grok 4.7, Doubao Seed 2.1 Pro/Turbo, Qwen 3.8 27B and Gemma 4 26B A4B IT. Doubao Pro selects the September 15 snapshot. Models and current prices
  • Copyable examples use the appropriate API. GPT tool workflows start with Responses; Claude 5.5 starts with native Messages. Model JSON exposes reviewed endpoint, reasoning and pricing-policy guidance. Fixed customer prices for the relevant GPT, Claude and Grok IDs disclose their cache-write and long-context ceilings rather than presenting them as maker base rates.
  • Five new guides in English and Chinese bring the library to 55 articles. They cover exact IDs, tool-history preservation, cost worksheets and task-based selection. MiMo's guide includes the announced October 21 V2.5 retirement. GPT guide · Claude migration · MiMo migration · GLM variants
  • Kimi now lists K2.6, K2.7 Code, K2.7 Code High-Speed and K3 with separate input, output and cached-input prices. The model pages link to exact maker documentation and show verified context limits. K3's fixed input rate covers uncached input and one-hour cache writes; cache hits use the published cache rate.
  • The quickstart and Gemini API guide clarify the Google Cloud chat route: use a KeepRouter key, a public Gemini ID and Chat Completions. Streaming is supported; Responses, Messages, Google-native generateContent and Live WebSockets are separate capabilities and are not exposed by this route.

What's new in KeepRouter. Have a request? Email support@keeprouter.com.

2026-09-29

More integration, troubleshooting and cost guides

  • The library now contains 50 guides, each in English and Chinese. New walkthroughs cover n8n, Dify, LangChain, LlamaIndex, Cline, Continue, Aider, Open WebUI and LibreChat, plus migration from Roo Code. Each states its setup and capability limits. Browse the guides
  • Reusable examples for difficult API decisions. Work through rate limits, interrupted streams, structured output, tool loops, context budgets, RAG and reasoning costs, document extraction, batch processing and local inference. Cost examples label their assumptions and link to current model prices. RAG cost worksheet · Document extraction evaluation

More Gemini models and clearer token usage

  • Nine Gemini chat models are available. Seven model IDs join the catalog. All nine listed IDs passed bounded text calls through KeepRouter. The new guide includes exact IDs, current price links, an OpenAI SDK example and the tested streaming, JSON and tool-conversation scope. Gemini API guide
  • Reasoning is counted once. Reported thinking tokens are included in completion usage, with the breakdown retained. Use each model page’s current input, output and cached-input rates to check cost. Models and prices
  • More precise model attribution. HappyHorse pages now identify Alibaba, their task and resolution variants, and official examples. Text Embedding 004 identifies Google and separates its Cloud lifecycle from Gemini API retirement. HappyHorse I2V · Text Embedding 004

2026-09-13

DeepSeek V4 Pro continues after September 14

  • The planned switchover is no longer happening. DeepSeek's current API notice says deepseek-v4-pro will continue after September 14, 2026, and that any later change will receive a separate notice. KeepRouter therefore keeps V4 Pro as a separately callable, separately priced model; no production route or customer rate was changed for this notice. DeepSeek's notice · V4 Pro pricing & specs

Veo 3.1, completion webhooks, and per-model clip lengths

  • Completion webhooks. Pass webhook_url when you submit a render and the gateway POSTs a signed event when the task finishes — no polling loop required. The event carries the same fields the poll returns, is signed with kr-signature (HMAC-SHA256 over <t>.<body>, using the secret returned once at submit), and is retried on a bounded schedule (about 1 min, 15 min, 1 h, 6 h) if your endpoint is down. Only public https destinations are accepted: internal, loopback and cloud-metadata addresses are refused before the task is ever submitted. The poll stays the source of truth and now reports the delivery state, so a broken endpoint is visible rather than silent. Video generation
  • Every poll response now carries expires_at. The clip URL and the gateway's tracking window share one 24-hour lifetime; the field makes that machine-readable instead of leaving you to compute it. Video generation
  • 4K and 1080p tiers added across the video catalogue. veo-3.1-fast-4k ($0.30/s), veo-3.1-4k ($0.60/s), and 1080p versions of happyhorse-1.0-i2v, happyhorse-1.0-r2v and happyhorse-1.0-video-edit ($0.24/s each). Check the selected model page for its current rate and supported resolution. Models & Pricing
  • Image-to-video, reference-to-video and video edit are published. happyhorse-1.0-i2v, happyhorse-1.0-r2v and happyhorse-1.0-video-edit join the text-to-video tier at $0.14/s (720p). They take their input through a media array — { "type": "first_frame" | "reference_image" | "video", "url": "https://…" } — where the URL must be publicly fetchable and a video input must be at least 3 seconds. Clip lengths are 3 to 15 seconds, refused locally when out of range. Video generation
  • The Veo 3.1 family is complete in both HD tiers. veo-3.1-lite (720p $0.05/s, 1080p $0.08/s), veo-3.1-fast (720p $0.10/s, 1080p $0.12/s) and veo-3.1 (720p and 1080p, $0.40/s). Completed test clips included synced audio. Models & Pricing
  • Veo 3.1 Lite is live on POST /v1/video/generations in two resolution tiers: veo-3.1-lite (720p) at $0.05/s and veo-3.1-lite-1080p at $0.08/s, each pinned so the request cannot be rendered above the tier it was priced at. At these customer rates, a 4-second clip costs $0.20 at 720p and $0.32 at 1080p. Models & Pricing
  • Clip lengths are checked per model before the render is submitted. Veo 3.1 accepts only 4, 6 or 8 seconds — it refuses anything else, and a min/max range would have let 5 and 7 through. The gateway now knows the difference between a range and a list, returns 400 invalid_duration with the lengths the model does accept, and, when you omit duration, sends the shortest length that model supports (4 seconds for Veo, 5 for the rest) so "omitted" can never mean "rejected".
  • Cached-input rate corrected for Kimi K2.7 Code (High-Speed). Cached input is now $1.90 / 1M tokens, matching the uncached input rate. The published rates are $1.90 / 1M in, $8 / 1M out, $1.90 / 1M cached. Uncached input and output rates are unchanged. Models & Pricing
  • Published prices receive regular checks. Pricing changes require review before publication. Check the exact model page for the current customer rate.

2026-09-12

Video generation, text-to-speech, and multimodal embeddings

  • Video generation is live on POST /v1/video/generations. Nineteen published model tiers across four families: Wan 2.7 (text-to-video, image-to-video, reference-to-video, and video edit, at 720p and 1080p), HappyHorse 1.0, PixVerse v3.5 through v6, and MiniMax H3. Video is an asynchronous task API — the submit call returns a task id and you collect the render from GET /v1/videos/{id} (or /v1/video/generations/{id}; both work). Billing is per second of clip, charged when the task is accepted. How the flow works
  • Video prices are explicit for each resolution tier. Each priced resolution has its own model ID: for example, wan2.7-t2v is 720p and wan2.7-t2v-1080p is 1080p. The gateway pins the selected tier so a request cannot be rendered at a higher resolution than the one priced. Models & Pricing is the current billing reference.
  • Tasks belong to the account that submitted them, and are tracked for 24 hours. A task's poll response carries a signed download URL, so it is readable only by the account that submitted it — another account gets the same 404 as an unknown id. After 24 hours the gateway reports 404 task_expired rather than relaying a state nobody can act on.
  • Text-to-speech is published. gemini-3.1-flash-tts-preview on POST /v1/audio/speech, billed per second of generated audio at $0.0005/s.
  • Multimodal embeddings are published. doubao-embedding-vision-250615 and -251215 embed text, images, and video on POST /v1/embeddings/multimodal. Use the dedicated multimodal endpoint shown on the model pages.
  • Prices now carry a unit. A video model is priced per second of clip and a speech model per second of audio, so /models, each model page, the Markdown twins, JSON-LD and /models.csv state the unit instead of rendering a per-second rate as a per-million-token one. /models.csv gained output_price_unit and output_price columns for the same reason.
  • New error codes for the video and audio routes: resolution_not_priced, invalid_task_id, task_not_found, task_expired, and upstream_account_unavailable (returned when an upstream account cannot serve a model, instead of forwarding the provider's own account message). Each is documented at /docs/errors.
  • The catalog now lists 80 models. The npm SDK experiment was removed: a gateway should be reachable with the client libraries you already use, so the OpenAI- and Anthropic-compatible routes remain the whole integration surface.

2026-09-11

DeepSeek V4.1 Flash and four more current API models

  • Five production-probed model IDs. KeepRouter now publishes deepseek-flash, gpt-6-astra, claude-fable-5-1, gemini-3.8-flash, and qwen3.8-max-0902 with maker-linked specifications and exact callable IDs.
  • Current fixed customer rates. Prices below are input / output / cached input per million tokens: DeepSeek Flash is $0.30 / $1.20 / $0.006 using the peak-hour rate as a fixed ceiling; GPT-6 Astra is $25 / $75 / $2, covering OpenAI's long-context cache-write ceiling; Claude Fable 5.1 is $20 / $50 / $0.25, covering the one-hour cache-write rate; Gemini 3.8 Flash is $0.75 / $3.75 / $0.075 through December 31, 2026; and Qwen3.8 Max 0902 is $2.50 / $6 / $0.25, covering cache writes. KeepRouter currently publishes one fixed input rate per model, so that same input rate applies to ordinary uncached input as well. Models & Pricing remains the billing source of truth.
  • DeepSeek migration stays explicit. The retired deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs remain temporary compatibility aliases to V4.1 Flash and now use the same fixed rate. DeepSeek has since confirmed that deepseek-v4-pro will continue after September 14 with its billing method unchanged, so V4 Pro remains separately callable and priced. Current notice
  • Kimi K3 route repaired. kimi-k3 now uses a production-probed route without changing its public model ID or customer price.
  • Retired Kimi K2.5 removed. Moonshot retired kimi-k2.5 on August 31, so it is no longer published or routed as an available model. Requests should move to a current, explicitly tested Kimi model instead of relying on a silent alias.
  • Restricted variants stay out of the public catalog. Claude Mythos 5.1 and Gemini 3.8 Flash Cyber require vetted-access programs, so KeepRouter does not present them as generally callable models.

2026-08-26

DeepSeek Vision Exp and refreshed V4 rates

  • New experimental multimodal model. deepseek-v4-flash-vision-exp accepts mixed text and image input and returns text. DeepSeek documents a 1M-token context window, a 384K maximum output, and support for Chat Completions, Messages, and Responses. Its historical model page now resolves to the current DeepSeek V4.1 Flash page.
  • V4 prices refreshed. V4 Flash, V4 Pro and Vision Exp use the fixed customer rates shown in Models & Pricing. KeepRouter does not switch these rates by time of day.
  • Callable ids stay explicit. Use deepseek-v4-flash, deepseek-v4-pro, or the experimental deepseek-v4-flash-vision-exp. Backend version labels are not public model ids.

2026-08-15

A bilingual AI gateway decision library

  • More useful comparisons. The comparison library now covers 15 buyer guides and product comparisons, including fit-based AI gateway and OpenRouter-alternative shortlists, managed versus self-hosted decisions, direct provider APIs, and seven additional gateway or cloud-platform comparisons. Shortlists are organized by operating fit rather than a synthetic ranking.
  • Answers and production playbooks. Eight new direct-answer pages explain routing, failover, latency, security, BYOK, gateway selection, and adoption timing. Three new editorial guides provide security, latency-measurement, and failover test plans.
  • Stable Simplified Chinese URLs. Every feature, audience, comparison, answer, and blog page now has a crawlable /zh/... HTML and Markdown twin with localized schema, reciprocal hreflang, and language-preserving internal links.
  • Clearer reading and discovery. Long-form pages now use answer-first verdicts, sticky contents, decision tables, checklists, step cards, evidence panels, FAQ blocks, and complete related-guide grids. Sitemap, llms-full.txt, AI index, RSS, SSR, and Markdown derive from the shared registries; root llms.txt enumerates every English page and links the five Chinese content hubs.

2026-08-13

DeepSeek V4 Pro is generally available

deepseek-v4-pro now resolves to DeepSeek's GA release DeepSeek-V4-Pro-0813. The callable model id has not changed. Its catalog metadata now records the maker-documented 1M-token context window, 384K maximum output, and August 13 release date.

  • The current stable DeepSeek ids are deepseek-v4-pro and deepseek-v4-flash. The retired

deepseek-chat and deepseek-reasoner aliases are no longer suggested by the admin preset.

  • DeepSeek now documents Chat Completions, Anthropic-compatible requests, and native Responses API

support for both V4 models. The admin includes separate Chat and Responses channel presets.

  • Both V4 models default to thinking mode at high effort; the maker API supports low, high, and max.

2026-07-26

New: DeepSeek V4 Flash — a 1M-token context window

deepseek-v4-flash, a DeepSeek model with a maker-documented 1M-token context window, is now in the production catalog. Its historical link now resolves to DeepSeek V4.1 Flash pricing & specs.

  • Also added: deepseek-v4-pro, same maker-documented 1M-token context, on the chat route.

Current rates for both are published in Models & Pricing.

2026-07-17

  • New: Kimi K3. kimi-k3 — Moonshot AI's flagship with a maker-documented 1M-token context window — is now on KeepRouter. See Kimi K3 pricing & specs.
  • More Kimi models. Added kimi-k2.7-code-highspeed (coding-focused, high-speed serving tier) and kimi-k2.5, joining the existing kimi-k2.6. See Models & Pricing.

2026-07-12

  • Public site redesigned. The homepage opens with a working request view backed by the production catalog and health probe. Models and prices use a product directory, and model pages pair facts with a copyable endpoint-specific request.
  • Documentation and trust pages reorganized. Quickstart, errors, security, company, comparison, policies, changelog, glossary, and SDK guides use a shared three-column shell with persistent navigation and page contents.
  • First-request path corrected. New examples default to free. Public model CTAs preserve the selected model through email verification, and API key enable and disable actions send boolean state.
  • Purchase facts moved before signup. The public site now states the minimum card top-up, processing fee formula, tax handling, refund boundary, request-log fields, cache default, and exact health-probe scope.

2026-07-10 · v0.5.0

  • Production operations complete. Production and staging are deployed separately; releases run through typecheck, Worker/frontend tests and build, then an isolated staging runtime/schema canary before production. Public staging keeps provider routing and self-serve OTP/email disabled by default rather than copying production credentials.
  • Safer operator sessions. The browser admin now uses a signed HttpOnly, Secure, SameSite=Strict cookie; the master admin token remains a non-browser break-glass path and is no longer stored by frontend JavaScript.
  • Hardened production sign-in. Transactional email delivery powers OTP, welcome and low-balance mail; a real bot challenge protects the production code-request flow. Staging's official test challenge is not treated as anti-bot protection.
  • Production billing. One-time Paddle card top-ups, customer billing portal, signature-verified webhooks, idempotent crediting, and refund/chargeback clawbacks are live.
  • Recovery path. Production D1 is exported on a schedule, with Worker rollback and D1 Time Travel documented for operators.

2026-07-05

  • New: a free model. free is priced at $0 per token. It is intended for testing, prototyping, and evaluating the API before moving to a paid model. See Models & Pricing.
  • Mistral AI models. Added mistral-large-latest and mistral-medium-latest; the codestral-latest code model and the devstral-medium-latest coding-agent model; the magistral-medium-latest and magistral-small-latest reasoning models; the ministral-3b-latest, ministral-8b-latest and ministral-14b-latest edge models; and the open-weight open-mistral-nemo.
  • Xiaomi MiMo & open-weight models. Added Xiaomi's mimo-v2.5 and mimo-v2.5-pro, plus the open-weight gpt-oss-120b (OpenAI) and gemma-4-31B-it (Google DeepMind). Current rates are published in Models & Pricing.

2026-07-04

  • Claude Fable 5 is back. claude-fable-5 is available again on KeepRouter. See Models & Pricing.

2026-06-26

  • Transparent top-up processing fee. Top-ups now show the selected credit amount and a separate card-processing fee before payment. Model usage is deducted at the published per-model rates, and the credited balance equals the amount selected.
  • Activity log. The Usage page now has a filterable request log (by model & status) with a per-request detail view and CSV export.
  • Billing history & invoices. A button on the Credits page opens the Paddle customer portal for invoices, receipts, and payment methods.
  • Quickstart guide at /docs/quickstart and an API error reference.
  • Low-balance email alerts send a one-time message after a paid balance crosses the warning threshold, plus a welcome email for new accounts.

2026-06-25

  • New: Security & Data Handling page. The /security page documents request-log fields, upstream forwarding, and optional response caching. Prompt and completion bodies are not written to the API request log; response caching is off by default.
  • Added Doubao models (ByteDance): doubao-2.0-pro, doubao-2.0-mini, doubao-1.5-pro, doubao-1.5-lite. See Models & Pricing.
  • Published per-page metadata and structured data, added year-long immutable asset caching, and reduced the page bundle size.

2026-06-24

  • Clearer image-model pricing. Image models now show a per-image price (e.g. $0.04 / image) instead of a token rate.
  • Full bilingual site. English and 简体中文 are selected from the browser language and can be changed manually.
  • New K-monogram logo and social preview cards.

2026-06-23

  • Service status page at /status.
  • Published Terms of Service, Privacy Policy, Acceptable Use Policy, and Refund Policy.

2026-06-22

  • Credit & debit card top-ups. Paddle processes prepaid, pay-as-you-go top-ups as Merchant of Record. There is no monthly subscription.