Engineering answer

What is the difference between an AI gateway and an API gateway?

An API gateway manages general service traffic through authentication, routing, rate limits, policy and observability. An AI gateway applies similar controls to model calls, then adds model-aware concerns such as canonical model IDs, provider selection, token usage, streaming events, tool-call continuity, multimodal payloads and AI-specific failover. Some products combine both layers; others place a focused AI gateway behind an existing API gateway.

Last reviewed 2026-08-26 · Editorial review: KeepRouter Editorial

The overlap is real, but the traffic is different

ConcernGeneral API gatewayAI gateway
Caller authenticationAPI keys, OAuth, mTLS or signed requestsGateway keys plus model and endpoint scope
Routing targetService, version, region or upstreamModel ID, provider route, endpoint family or fallback
Usage unitRequests, bytes and durationRequests plus input, output, cached tokens or generated assets
StreamingGeneric HTTP or event streamsModel-specific SSE events, partial output and terminal usage
PolicyRate, identity, network and schema rulesModel allowlists, spend limits, prompt or response controls
Failure handlingRetry service instancesRetry or fail over only within model and side-effect boundaries

A standard gateway can proxy model HTTP traffic, but it does not automatically understand token accounting, tool-call state, model capability, provider-specific errors or the difference between Chat Completions and Responses. An AI gateway can implement those concerns, although the exact feature set varies by product.

Three common architectures

  1. Focused AI gateway only. A small team points its SDK at a managed model gateway and lets that service own model access and billing.
  2. API gateway in front of an AI gateway. The outer layer handles company identity, network policy and tenant admission. The inner layer handles models, tokens and provider routes.
  3. AI plugins inside the existing gateway. Platform teams extend Kong, Envoy or another gateway with model-aware plugins and keep the data plane inside their infrastructure.

The right arrangement depends on who already owns traffic policy. If an API platform team has mature identity and network controls, replacing them with a separate public entrypoint may create duplicate policy. If a small product team has no gateway platform, operating one solely for model calls can be unnecessary work. See when an AI gateway is useful before adding another service.

Follow the request through both layers

Client -> API Gateway -> Application -> AI Gateway -> Model provider

HopPrimary responsibilityIdentifier that should continue
Client to API GatewayUser session, tenant admission and public rate policyTrace ID or request ID
API Gateway to applicationService routing, identity claims and network policyTrace ID plus tenant ID
Application to AI GatewayTask context, approved model and product limitsTrace ID plus AI request ID
AI Gateway to providerModel route, provider credential, usage and fallbackAI request ID plus provider request ID

The application remains a useful boundary between the two gateways. It can translate an authenticated user request into a narrowly scoped model request without exposing the model gateway key to the client.

Deployment choice matrix

Starting pointUsually simplestCheck before shipping
No existing gateway platformUse one focused AI gateway for server-side model callsKey scope, endpoint support, logs, spend limits and rollback
Mature API gateway and platform teamKeep the API gateway at the edge and place the AI gateway behind the applicationTimeout, streaming, body limits, identity propagation and duplicate limits
Required private data plane or gateway standardAdd model-aware plugins or self-host the AI layerProvider compatibility, operational ownership, upgrades and evidence export
One provider and one stable modelStart with the provider SDK, then add a gateway when a specific control is neededAvoid an extra layer without a measured requirement

A short decision tree is enough: if enterprise identity and network policy already live at the edge, keep them there. If model routing, usage, or failover needs a separate owner, add the AI layer behind the application. If neither condition applies, document the requirement before adding another hop. The managed versus self-hosted comparison helps resolve the data-plane choice.

Questions to settle before combining the layers

Write down which layer authenticates the end user, which key reaches the model gateway, where tenant and feature identifiers are attached, and which system is authoritative for rate and spend limits. Decide where prompt content may be logged and how request IDs cross both layers. Two dashboards with different request IDs will slow incident response.

Also verify timeouts. A general API gateway may close a connection before a long model response or streaming session finishes. Request and response size limits can reject image blocks or long contexts even when the model route accepts them. The AI gateway security answer covers trust boundaries, while the observability feature describes the KeepRouter evidence surface.

Correlate one request, not two dashboards

LayerExample record
API Gatewaytrace_id=req_7f2, tenant_id=team_42, status=200
Applicationtrace_id=req_7f2, ai_request_id=ai_91c, feature=answer
AI Gatewayai_request_id=ai_91c, model=approved-model-a, fallback=false, usage_total_tokens=1842
Provider adapterai_request_id=ai_91c, provider_request_id=provider_abc, attempt=1

Propagate identifiers in headers or structured context, then record them as fields rather than embedding them only in free-form messages. Do not put prompts, secrets, or raw authorization headers into correlation fields. During a test, begin with the client trace ID and confirm that you can find the final model, usage, charge, latency, and provider result. Use the gateway latency answer to separate the two gateway hops from model generation time.

Keep the contract narrow

Do not advertise a single universal AI endpoint if the implementation supports only one operation. Document Chat Completions, Responses, Messages, embeddings, images and audio separately. A useful gateway makes those boundaries easier to inspect. It should not hide them behind a generic proxy label.

Frequently asked questions

Can an ordinary API gateway proxy LLM calls?

Yes, but generic proxying does not add model catalogs, token accounting, provider-aware routing or model-specific stream handling by itself.

Do I need both gateways?

Only if the layers have distinct owners and policies. Many small teams need one focused layer; larger platforms may keep enterprise traffic controls in front.

Which layer should enforce spend limits?

Choose one authoritative layer and reconcile its model usage evidence. Duplicate independent limits can produce confusing failures.

Can I put an AI gateway behind Kong or Cloudflare?

Often yes, subject to timeout, streaming, body-size, identity and logging configuration on both layers.

Is an AI gateway only for LLM text?

No. Some gateways support images, embeddings, audio and other modalities, but each operation needs an explicit route and capability check.

Sources reviewed

  1. [1] Kong AI Gateway
  2. [2] Cloudflare AI Gateway
  3. [3] Envoy AI Gateway documentation

Related guides

Test the contract with a real model

Create a narrowly scoped key, select a model from the live catalog, and run the exact request shape your application depends on.

Create a free key · View live models and pricing · Read as Markdown