An AI gateway is an intermediary on the request path between applications and model services. It can centralize authentication, policy, routing, usage accounting and traffic controls. Its data plane handles requests and responses; its control plane manages configuration such as provider credentials, routes and budgets. Calling the whole gateway only a control plane misses the request-serving responsibility.
Model routing selects a model or endpoint for a task. Load balancing distributes work among eligible endpoints. Fallback selects another path after a failure. These mechanisms can work together, but they solve different problems.
When the extra layer earns its cost
A small application can begin with a direct SDK behind a narrow application interface. A shared gateway becomes useful when several applications duplicate access, accounting, quota or routing logic, or when centralized policy is a concrete requirement.
The number of providers is not a universal adoption threshold. One provider serving many teams may justify a gateway. Two providers serving a small prototype may not. The added component needs availability, capacity and an operating owner.
| Responsibility | Gateway can centralize | Still needs application context |
|---|---|---|
| Authentication | Verify an application/virtual key | End-user identity and resource authorization |
| Routing | Enforce eligible models/endpoints | Required task capabilities and quality |
| Traffic | Rate/concurrency limits, deadlines, retries | Whether delayed or degraded output is acceptable |
| Accounting | Attempts, usage and rate references | What counts as a successful user outcome |
| Caching | Store/reuse eligible responses | Access, freshness, policy and semantic validity |
| Filtering | Configured input/output controls | Domain rules and authorization at actual tool execution |
A common API shape reduces adapter work; it does not guarantee identical tool calls, structured output, tokenization, streaming, reasoning or error semantics across providers.
Interview design: shared access for the learning platform
Assume lesson feedback, quiz explanation and offline content checks share model access.
Functional requirements
- Authenticate each application and propagate trusted account/workload scope.
- Select only endpoints approved for the request's capabilities and data policy.
- Apply request, token, concurrency and spend controls.
- Stream or return results while recording all attempts and their outcomes.
- Support bounded fallback and versioned routing changes.
Non-functional requirements
- Prevent clients from overriding mandatory provider/access restrictions.
- Bound gateway overhead and preserve end-to-end deadlines.
- Avoid retry storms and cross-account cache leakage.
- Keep accounting and request correlation durable enough for reconciliation.
- Remain available through a gateway-replica failure and define behavior when shared policy services fail.
Start with explicit routes: quiz explanations and feedback use evaluated configurations; offline jobs use a separate quota class. Add learned routing only if its quality/cost benefit exceeds the extra inference, evaluation and operating complexity.
Read diagram source
flowchart TD
A[Application and trusted request scope] --> G[Gateway replicas]
P[Versioned policy and provider registry] --> G
G --> E[Eligibility: access, data, capabilities]
E --> B[Reserve quota, concurrency and budget]
B --> C{Eligible cached result?}
C -->|Yes| O[Return with provenance]
C -->|No| R[Choose eligible endpoint]
R --> M[Provider or self-hosted model]
M --> H{Outcome}
H -->|Success| O
H -->|Retryable and budget remains| R
H -->|Final or uncertain| F[Explicit failure or recovery path]
O --> U[Reconcile usage and release reservations]
F --> U
The actual ordering of cache lookup and reservation depends on which work is billed and rate-limited. Authentication, access and cache eligibility must happen before a cached answer is exposed. Every fallback repeats the required eligibility checks.
Eligibility first, optimization second
A route decision should first exclude endpoints that violate a hard requirement:
- Required model modality, context capacity, tool or output-schema support.
- Permitted data processing location and provider/data-retention policy.
- Account access and organizational model restrictions.
- Available deadline, quota and budget.
- Required evaluated task quality.
Then rank eligible choices by the product's objective. A cheap endpoint outside the permitted region is not a valid candidate. If none qualify, return a clear unavailable result or an approved degraded mode; do not silently relax constraints.
| Routing approach | Useful when | Main tradeoff |
|---|---|---|
| Static/task rule | Task classes and requirements are known | Simple and explainable; rules need maintenance |
| Cost-aware | Several choices meet quality/capability requirements | Must include retries, evaluation and output-length effects |
| Latency/load-aware | Endpoint conditions vary | Estimates can be stale; avoid sending everyone to the same endpoint |
| Semantic/learned router | Task difficulty or type is hard to describe with rules | Router inference, errors and distribution drift |
| LLM classifier | Classification needs model judgment | Adds a model call, cost and another failure path |
| Cascade | A first result can be checked and escalated | Sequential latency and payment for both attempts |
RouteLLM studies learned routing between stronger and weaker models. Its reported results belong to particular models and evaluations. Routing before generation differs from a cascade that generates a cheaper answer and then decides whether to escalate. Neither inherently supplies outage recovery.
Calculate cascade economics
With hypothetical costs of $0.002 for a first model, $0.020 for an escalated model, and $0.001 for a verifier per task:
expected cost = first call + verification + escalation rate × second call
at 20% escalation: 0.002 + 0.001 + 0.20 × 0.020 = $0.007
at 90% escalation: 0.002 + 0.001 + 0.90 × 0.020 = $0.021
The second case costs more than using the stronger model once, before adding routing overhead. Compare accepted quality and latency as well as dollars. Model-reported confidence is not automatically calibrated; a weak verifier can accept precisely the answers that needed escalation.
Retry errors by meaning, not only status family
| Outcome | Typical response |
|---|---|
| Temporary capacity/rate limit | Respect provider guidance, back off or use an independently eligible endpoint |
| Transient service/network error | Bounded retry if the request is safe to repeat and time remains |
| Invalid schema/context too long | Correct or reject the request; repeated identical calls usually do not help |
| Authentication/authorization failure | Fix configuration or deny access; do not bypass the restriction |
| Unknown model/deployment | Treat as configuration/lifecycle issue unless a tested mapping explicitly handles it |
| Safety/policy refusal | Follow application policy; do not route around the restriction merely to obtain an answer |
| Timeout after a possible external effect | Reconcile the effect before repeating it |
429 is itself a 4xx status, so “never retry any 4xx” is too broad. Conversely, not every 5xx warrants another attempt beyond the deadline. Preserve meaningful error details without exposing credentials or private content.
Use an end-to-end attempt/time budget, exponential backoff with jitter where appropriate, and Retry-After guidance when supplied. Avoid multiplying retries at every layer. One application retry around three gateway attempts can already make six provider attempts.
A circuit breaker stops sending normal traffic to a failing dependency while it recovers, then probes cautiously. Scope health to the relevant endpoint/account/limit: one tenant's invalid key or quota exhaustion need not disable service for everyone.
More keys do not automatically create more quota. Limits may be shared at account, organization, region or deployment level. Plan legitimate capacity; do not treat key rotation as a way to bypass provider limits.
Streaming changes fallback behavior
Before output reaches the client, a safe generation-only request may be retried under policy. After a partial answer is emitted, appending a different model's fresh answer can create contradictory text or malformed tool/JSON output. End the stream with an explicit failure, or use a designed restart/resume protocol that tells the client what to replace.
A tool action may have succeeded even if its response was lost. A new model or provider does not make that action safe to repeat. Use operation IDs and effect reconciliation.
Current product options and what to inspect
Documentation reviewed in September 2026:
| Option | Relevant surface | Adoption check |
|---|---|---|
| LiteLLM | Router/proxy controls and provider adapters | Retry ownership, caller overrides, supported error semantics and shared-state operations |
| OpenRouter | Managed provider selection and restrictions | Required-parameter support and actual provider/data-policy eligibility |
| Cloudflare AI Gateway | Managed gateway, caching, rate controls and routing features | Feature maturity, including beta spend-limit/dynamic-routing surfaces |
| Portkey | Gateway routing and operational controls | Hosting mode, policy scope, data paths and contractual limits |
| Kong AI Gateway | AI capabilities within an API gateway platform | Required plugins, editions, supported providers and operating fit |
| Agent Router | Current destination of Envoy AI Gateway's provider-fallback documentation | Supported routing/retry behavior and deployment integration |
For example, OpenRouter documents require_parameters for excluding providers that do not support all requested parameters. Do not assume every adapter refuses unsupported options by default. LiteLLM documents separate retry configuration and caller overrides; enforce product limits at a boundary the caller cannot weaken. Verify the installed version rather than copying an old comparison table.
Managed options reduce infrastructure work but still require policy, cost and incident ownership. Self-hosting the proxy does not keep model payloads local when the next hop is an external provider.
Availability, budgets and observability
Run appropriate redundant gateway replicas and avoid keeping authoritative budgets only in process memory. Reserve estimated spend atomically before concurrent work; reconcile all attempts afterward. Decide whether a missing budget/policy dependency means fail closed, a restricted preallocated allowance, or another explicit behavior.
An alerting dashboard is not a hard spend cap. Provider usage may arrive late, and cancellation does not guarantee zero further charges. See FinOps controls.
Record request/task ID, account/workload scope, route-policy revision, attempted endpoints, outcome, latency, token usage and estimated/reconciled cost. Redact before exporting content. Use bounded dimensions for aggregate metrics; keep per-request identifiers in traces.
| Observed problem | Repair | Cost/benefit |
|---|---|---|
| Every retry hits the same exhausted quota | Model quota domains and cool down the affected scope | More accurate availability state |
| Alternate provider violates required behavior | Capability contracts and paired evaluations | Fewer eligible fallbacks |
| Cache hit leaks private context | Scope by access, relevant versions and freshness | Reduced hit rate but correct results |
| Gateway is a new single point of failure | Redundant data plane and tested shared dependencies | Extra operating cost |
| Latency router creates a traffic stampede | Load-aware selection and controlled exploration | More routing complexity |
| Retried jobs exceed the tenant budget | Shared reservation and total-attempt accounting | Persistent state and reconciliation |
Interview questions and answer checks
- Is a gateway a control plane? It usually includes control-plane configuration and a data plane that serves requests.
- Can model routing be useful with one provider? Yes, if different models or deployments serve different requirements; centralized access/accounting may also help.
- Why can a fallback hurt reliability? It may lack capacity, change behavior, violate policy or repeat an uncertain effect.
- Do multiple API keys multiply quota? Not necessarily; inspect the provider's quota scope.
- Why not retry every error? Permanent/configuration failures waste capacity, and uncertain writes can be duplicated.
- What must happen before semantic cache reuse? Check access, task suitability, freshness and relevant versions; similarity alone is insufficient.
- How do you validate a learned router? Compare end-to-end quality, cost, latency and important slices against a simple baseline; monitor distribution drift.
- When is a gateway overkill? When the application can implement the required boundaries clearly with less operational complexity; provider count alone is not decisive.
Final notes
Remember eligible → selected → attempted → accounted. Preserve policy through every route and retry. Centralize controls when doing so makes the system easier to operate, and prove that the new shared layer can meet the reliability requirements it inherits.