Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an AI Gateway and Model-Routing Service

By Anup Rai8 min readReviewed September 2026

Interview problem: several product teams use different model providers. Create one governed API that enforces identity, budget, model eligibility, regional processing constraints and reliable usage accounting.

This is an illustrative design. Targets and prices are assumptions. A gateway can enforce an application policy; it cannot make every upstream model behave identically.

1. Scope and requirements

Clarify whether the gateway handles text only, tools, streaming, images, batches or embeddings. Assume text generation with optional structured output, two approved providers and one internal model. Tools execute in the caller's application, not in the gateway.

Functional requirements

  1. Authenticate applications and derive their tenant, use case and policy.
  2. Route requests only to eligible models and processing regions.
  3. Enforce request, token and spending limits before dispatch.
  4. Normalize common response fields while retaining provider-specific capabilities explicitly.
  5. Support streaming, cancellation, bounded retries and usage attribution.
  6. Version routes and policies, audit changes and roll back a rollout.

Non-functional requirements

  1. Add less than 50 ms p95 gateway overhead for ordinary admitted requests, excluding upstream generation.
  2. Target 99.95% monthly gateway availability; report end-to-end model availability separately.
  3. Prevent unauthorized fallback or cross-tenant cache reuse.
  4. Bound financial exposure during concurrent requests and uncertain provider outcomes.
  5. Keep raw prompts out of routine telemetry; redact and restrict sampled debugging records.

2. Estimate the workload

At 500 requests/s, mean 2,000 input and 400 output tokens, the gateway forwards 1M input tokens/s and eventually 200,000 output tokens/s at steady state. A ten-second mean request duration implies about 5,000 concurrent upstream requests. The stream relay must handle many open connections; CPU utilization alone is not the capacity signal.

Assume an eligible route's hypothetical price is $2/M input and $10/M output. A typical call costs .004 + .004 = $0.008; 500/s sustained for an hour is about $14,400. This is why budget enforcement cannot wait for an end-of-day dashboard. The rate is a peak assumption, not a forecast of monthly spend.

3. Baseline and its flaws

Architecture / visual model
flowchart LR A[Product applications] --> G[Authenticated proxy] G --> P[One model provider] P --> G G --> L[(Usage log)]
Read diagram source
flowchart LR
 A[Product applications] --> G[Authenticated proxy]
 G --> P[One model provider]
 P --> G
 G --> L[(Usage log)]

Start with a fixed allowed route and schema validation. The first experiment should establish pass-through overhead, error semantics and accounting accuracy.

Failure Repair Benefit Added cost
Products retry while proxy also retries One attempt budget across layers Contains amplification Shared deadline/attempt contract
Provider outage triggers illegal region fallback Filter eligibility before ranking Preserves processing rules Some requests must fail
Concurrent requests overspend one quota Atomic reservation before dispatch Bounded committed exposure Reservation store and settlement
Streams end without usage metadata Mark uncertain usage and reconcile Honest accounting Temporary budget holds
Common API hides unsupported features Capability negotiation and validation Predictable semantics Less universal abstraction

4. Detailed architecture

Architecture / visual model
flowchart TD APP[Applications] --> AUTH[Identity and request validation] AUTH --> POL[Policy and capability filter] POL --> RES[Atomic budget reservation] RES --> ROUTE[Deadline-aware route choice] ROUTE --> A[Provider A adapter] ROUTE --> B[Provider B adapter] ROUTE --> C[Internal serving adapter] A --> STREAM[Stream normalization and cancellation] B --> STREAM C --> STREAM STREAM --> APP A --> EVT[Append usage and attempt events] B --> EVT C --> EVT EVT --> SETTLE[Idempotent settlement and reconciliation] SETTLE --> LEDGER[(Budget ledger)] LEDGER --> RES ADMIN[Approved policy changes] --> CFG[(Versioned routing snapshot)] CFG --> POL CFG --> ROUTE EVT --> OBS[Quality, latency, cost and error telemetry]
Read diagram source
flowchart TD
 APP[Applications] --> AUTH[Identity and request validation]
 AUTH --> POL[Policy and capability filter]
 POL --> RES[Atomic budget reservation]
 RES --> ROUTE[Deadline-aware route choice]
 ROUTE --> A[Provider A adapter]
 ROUTE --> B[Provider B adapter]
 ROUTE --> C[Internal serving adapter]
 A --> STREAM[Stream normalization and cancellation]
 B --> STREAM
 C --> STREAM
 STREAM --> APP
 A --> EVT[Append usage and attempt events]
 B --> EVT
 C --> EVT
 EVT --> SETTLE[Idempotent settlement and reconciliation]
 SETTLE --> LEDGER[(Budget ledger)]
 LEDGER --> RES
 ADMIN[Approved policy changes] --> CFG[(Versioned routing snapshot)]
 CFG --> POL
 CFG --> ROUTE
 EVT --> OBS[Quality, latency, cost and error telemetry]

The gateway's data plane and control plane have different failure requirements. An operator dashboard outage should not stop all generation; revocation of a compromised application must still take effect within the agreed bound.

5. API and records

POST /generations includes an application-visible model alias, messages, supported response schema, output cap, idempotency key and deadline. The caller cannot choose an arbitrary upstream URL or set its own tenant identifier.

Entity Stored fields Why
Policy snapshot ID, tenant, model/region allowlist, expiry, limits Reproduce an eligibility decision
Attempt request ID, attempt number, provider, model revision, start/end, outcome Separate logical work from paid calls
Reservation tenant, request ID, maximum exposure, state, expiry Prevent parallel overspend
Settlement attempt ID, measured input/output, charged units, evidence Deduplicate usage and reconcile bills

A reservation must cover the permitted retry/candidate budget, not merely the cheapest first attempt. Where providers expose different tokenizers, estimate conservatively using the selected route and settle measured usage. A hard spend guarantee requires contracts for unknown or delayed charges; state the bounded uncertainty instead of claiming exact real-time billing.

6. Route and execute

  1. Authenticate and determine permitted use-case policy.
  2. Reject unsupported modalities, schemas, token limits or destinations.
  3. Reserve the maximum allowed attempt budget atomically against the tenant allowance.
  4. Rank only eligible routes by a tested quality/latency/cost policy.
  5. Pin the route snapshot and dispatch with the original deadline.
  6. Relay events with backpressure. Retain the actual model identity and finish reason.
  7. Settle known usage once, hold uncertain exposure and reconcile provider records.

A fallback is a new model behavior. Test its quality, tool-call shape and refusal/error behavior. Do not silently label its answer as the originally requested model's output.

7. Retries and degradation

Situation Behavior
Rejected before dispatch Release reservation; no inference retry needed
Rate limited before any output Retry eligible route only within deadline and reserved budget; use jitter
Timeout with unknown completion Record uncertainty; avoid assuming zero charge
Partial stream delivered Explicit interruption or retained-event resume; no invisible replacement answer
Budget store unavailable Fail closed for spend-controlled traffic or use a separately approved bounded allowance
Provider quality regresses Disable route by evaluated signal; preserve a known approved version where available

A circuit breaker limits repeated failing calls. It is not a quality detector; monitor product-specific evaluation as well as HTTP status.

8. Cost and benefit

For an illustrative one million successful tasks per month, lowering average complete-task cost from $0.014 to $0.010 saves $4,000 before gateway operating cost. If the gateway costs $3,000 monthly and adds $2,000 of allocated maintenance, the change loses $1,000 on that accounting boundary. Governance or reliability may still justify it; call that benefit explicitly.

Decision Benefit Tradeoff
Static routing first Easy to audit and reproduce Less adaptive optimization
Learned routing later Can match difficulty to cost Router evaluation, drift and failure modes
Exact response cache Low-cost repeated safe requests Invalidation and identity/version keys
Semantic cache More reuse False-equivalence risk; needs separate evaluation
Multi-region gateway Better local resilience Policy propagation and budget consistency

9. Release and interview follow-ups

Test concurrent quota depletion, duplicate settlement, provider timeouts, a revocation during a stream, unavailable policy storage and an unsupported schema. Canary one application; compare complete-task quality, total cost and p95 overhead before expanding.

Q1: Why can the cheapest model be an expensive route?

Sample answer: It may need more retries, longer outputs, human repair or escalation. I compare cost per acceptable task, including routing and fallback calls, against a fixed quality and latency requirement.

Q2: Can provider compatibility remove vendor differences?

Sample answer: It can normalize a useful subset. It cannot guarantee identical tokenization, tool-call semantics, structured-output support, safety behavior, context limits or billing. The contract should expose unsupported capabilities explicitly.

Q3: What must an outage fallback preserve?

Sample answer: Authorization, regional/data-processing restrictions, required capabilities, budget and a tested minimum quality. If no route satisfies those constraints, fail explicitly instead of sending data to an unapproved service.

Closing remarks

I would ship a thin governed gateway with static approved routes, atomic reservations, explicit streaming semantics and reconciled usage. Adaptive routing comes after a reliable evaluation loop. The key tradeoff is centralized control versus an additional critical dependency.

Recall Explain
Filter before rank Policy determines eligibility
Reserve before send Concurrency can overspend
Count attempts and tasks A retry is still work and possibly cost
Fail explicitly An ineligible fallback is not resilience

Tip: Separate gateway overhead from upstream generation latency. Combining them hides which system must improve.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an Adaptive AI Learning Tutor
NEXT LESSONDesign a Deadline-Aware Batch Inference Platform →

Explore the diagram