System designby Learnastra

Speak before revealing

Interview question bank

100+ interview questions with worked answers and follow-ups. Select a topic, try your answer, then compare the reasoning. Checking a box records your own assessment; it does not grade your answer.

Foundation · Question 1System design interview framework

What is system design, and how would you begin “design file sharing”?

Reveal a model answer

“System design defines the components, stored data, interfaces and interactions needed to meet requirements. For file sharing I first ask who uploads, who downloads, size limits, and whether links are public, private, expiring or revocable. Then I agree on volume and what success means, and trace one upload before adding capacity.”

What the answer must demonstrate: Connect each clarification to an architectural consequence.

Foundation · Question 2System design interview framework

What is a correctness invariant? Give one for an upload-and-download API.

Reveal a model answer

“An invariant is a condition the system must preserve. Here, an incomplete upload must never become downloadable. I represent upload state explicitly and allow downloads only after completion is verified. The download handler can enforce this rule by checking the stored upload state before serving the file.”

What the answer must demonstrate: Give an enforceable rule, not an adjective.

Why start with one application and database?

Reveal a model answer

“It makes the complete request and stored state understandable. I can show which changes must succeed together, verify the rules that keep the data correct, and measure capacity. I split or replicate components when a workload, reliability requirement, or ownership boundary creates a reason, rather than assuming that a distributed diagram is inherently better.”

What the answer must demonstrate: Logical clarity should survive changes in physical scale.

A file service averages 2.3 downloads/s but may peak at 1,000/s. Is average QPS enough to choose one server?

Reveal a model answer

“Not from that average alone. I need peak request rate, average and large-file sizes, connection duration, and a per-server load test at the target latency. Traffic can be concentrated into short bursts. I would state the peak assumption and size for it, including a server failure.”

What the answer must demonstrate: Do not equate average QPS with capacity.

Why define a data model before naming a database product?

Reveal a model answer

“The model tells me what must be stored together and which queries must be efficient. For file sharing I need ownership, upload state, expiry and the public-token mapping checked by identifier. A transactional metadata database can enforce those relationships; I evaluate products after deciding durability, throughput and failure requirements.”

What the answer must demonstrate: Explain access patterns and constraints.

An upload completion commits but its response is lost. How should a retry behave?

Reveal a model answer

“A timeout means the client does not know the outcome. I keep a stable upload identifier and make the completion operation inspect its existing state. Retrying completion for an already-ready upload returns the same worksheet. I recover the existing outcome before creating a new upload.”

What the answer must demonstrate: Explain what happens if the server saves the result but the response is lost.

How do you answer “Why a CDN?” without a buzzword list?

Reveal a model answer

“Repeated downloads request identical bytes. A CDN can reduce origin traffic and serve a nearby copy. I would use versioned public objects where possible. If a link is private or revocable, I must define the authorization and cache lifetime so an old edge copy cannot bypass the promised access policy.”

What the answer must demonstrate: Name the benefit and the access-policy cost.

How do you close the interview?

Reveal a model answer

“I would recap the agreed user actions, trace the main path briefly, and state the key choices: durable upload states, independent byte delivery, and retryable completion. Then I would identify the first measured scaling limit and one remaining risk, such as revocation latency, with how I would test it.”

What the answer must demonstrate: Summarize decisions and limits rather than reciting components.

Does growing traffic mean a modular monolith must become microservices?

Reveal a model answer

“No. I can run multiple instances of the same application when its durable state is shared appropriately. I would extract a capability when independent capacity, releases or ownership justify the extra coordination. In checkout, keeping orders and stock reservations together preserves a useful local transaction; catalog search can scale separately as a derived view.”

What the answer must demonstrate: Distinguish server count, deployment boundaries and transaction boundaries.

Foundation · Question 1HTTP APIs and request lifecycle

What is an API, and what happens in a GET /orders/O17 request?

Reveal a model answer

An API is a contract between programs for an operation and its inputs, results, and failures. For GET /orders/O17, the client resolves the service name and establishes or reuses a protected connection. The edge routes the request; the order service validates the credential, derives user U9, checks U9’s permission for O17, and returns an authorized representation.

What the answer must demonstrate: Define the contract before tracing the complete request path.

Foundation · Question 2HTTP APIs and request lifecycle

Why use TLS if the API already checks a token?

Reveal a model answer

“The token identifies or authorizes the caller, but a plaintext network could expose or alter it. TLS protects the communication and authenticates the server endpoint. I still validate the token and resource permission inside the service.”

What the answer must demonstrate: Separate transport protection from access checks.

Applied · Question 3HTTP APIs and request lifecycle

Does HTTP/3 eliminate head-of-line blocking everywhere?

Reveal a model answer

“No. QUIC avoids TCP’s cross-stream loss-delivery blockage, but each stream still has ordering requirements and the application, queues, or shared resources can block progress. I choose it for actual transport needs, not as a blanket latency guarantee.”

What the answer must demonstrate: Name the specific bottleneck that changes.

Applied · Question 4HTTP APIs and request lifecycle

How are safe and idempotent requests different?

Reveal a model answer

“Safe methods do not ask for a state-changing action. Idempotent methods have the same intended effect when repeated. A deletion can be idempotent while still changing state. I do not use a GET to trigger a purchase merely because it is easy to call.”

What the answer must demonstrate: Explain intended effect, not identical response bytes.

Applied · Question 5HTTP APIs and request lifecycle

What does 202 Accepted tell the caller?

Reveal a model answer

The request was accepted for processing, not completed. For our recoverable API, I durably commit an operation record and outgoing intent before 202, then return an operation ID and status location. HTTP 202 alone does not establish that storage guarantee.

What the answer must demonstrate: Distinguish acceptance and completion.

Applied · Question 6HTTP APIs and request lifecycle

Would you choose REST or gRPC for every service?

Reveal a model answer

“I choose from client compatibility, schema tooling, and streaming needs. A public browser-facing API may use resource-oriented HTTP/JSON; internal typed calls may use gRPC. Both still need deadlines, authorization, and a defined retry contract.”

What the answer must demonstrate: Avoid assigning storage guarantees to a protocol.

Applied · Question 7HTTP APIs and request lifecycle

What belongs in a cursor for order history?

Reveal a model answer

“A stable position in the chosen order, such as the last creation timestamp plus a unique order ID. The service validates it, applies the same ordering, and caps page size. I also define whether new or deleted records can change later pages.”

What the answer must demonstrate: Match the cursor to the index and contract.

Applied · Question 8HTTP APIs and request lifecycle

How do you rename a required response field safely?

Reveal a model answer

“I cannot assume all clients update together. I might serve both fields during migration or introduce a versioned contract, measure adoption, and retire the old field under an explicit policy. I test mixed client/server versions.”

What the answer must demonstrate: Describe a mixed-version rollout.

What is capacity estimation? Estimate QPS for one million users making ten requests a day.

Reveal a model answer

“Capacity estimation converts a workload into rates and resource needs. Here one million users × ten requests is ten million requests/day. Dividing by 86,400 seconds gives about 116 requests/s average. I still need peak concentration, bytes per request, latency targets and failure headroom before choosing server capacity.”

What the answer must demonstrate: Show denominator and units.

Can a system have high throughput and high latency?

Reveal a model answer

“Yes. A batch worker may finish thousands of items a second while each item waits minutes in a queue. Throughput describes the completion rate; latency measures one item’s elapsed time. I would measure queue wait and processing time separately.”

What the answer must demonstrate: Distinguish work rate from wait time.

How much storage do 200 GB/day of uploads need after a year?

Reveal a model answer

“Without deletion, 200 × 365 is 73,000 GB, or 73 TB in decimal units. That is logical originals. I would separately add derived images, indexes, copies, and backups, then apply the retention policy.”

What the answer must demonstrate: Separate logical data from physical overhead.

What happens at 2,000 requests/s if average latency grows from 50 to 500 ms?

Reveal a model answer

“Assuming both measurements cover the same system in stable operation, the average number of requests in progress grows from about 100 to 1,000. That can exhaust memory or connection pools even without a traffic increase. I would inspect downstream latency and bound admitted work.”

What the answer must demonstrate: Use seconds and matching averages.

Should a cache hold 20% of yesterday’s requests?

Reveal a model answer

“Requests are not stored objects. I estimate distinct hot keys and bytes per entry. If a million requests hit one record, that is one cache entry. I use observed reuse and eviction behavior to choose the working set, then account for replication and overhead.”

What the answer must demonstrate: Count distinct retained entries.

Three servers can just meet peak. Is that a resilient design?

Reveal a model answer

“Not if the requirement includes surviving a server failure at that peak. I calculate the remaining capacity after the failure and keep headroom for imbalance. If two survivors cannot meet the objective, I add capacity, reduce admitted work, or agree on degraded behavior.”

What the answer must demonstrate: Calculate surviving capacity.

An API runs five database queries. Which QPS matters?

Reveal a model answer

“Count both. At 2,315 API requests/s and five queries per request, the database receives about 11,575 operations/s before retries or cache effects. I would check whether each query is needed, indexed and independent of the others.”

What the answer must demonstrate: Explain amplification rather than hiding it.

A photo service serves 2,315 peak views/s at 100 KB each. Which measurements would change the storage or delivery design?

Reveal a model answer

“The stated peak is 2,315 × 100 KB = 231.5 MB/s, about 1.85 Gb/s before overhead. I would measure repeated-key reuse and permission constraints to evaluate a CDN, and measure metadata and CPU costs separately to decide where scaling helps. Peak QPS alone cannot determine daily delivered bytes or metadata growth; those require daily volume and stored bytes per upload.”

What the answer must demonstrate: Use a number to justify a decision.

What is the difference between reliability and availability?

Reveal a model answer

“Availability asks whether an eligible checkout operation can complete under its success definition. Reliability asks whether the service performs its specified function correctly over time and under promised conditions. A reachable system that double-charges an order is incorrect; that purchase must also count as unsuccessful in an end-to-end availability measure. I define the outcome and measurement window rather than treating reachability as either guarantee.”

What the answer must demonstrate: Use the same example for both qualities.

When would you scale vertically before sharding?

Reveal a model answer

“If the database fits on one larger instance and measured CPU, memory, or I/O is the bottleneck, vertical scaling can buy capacity with a smaller operational change. I would also keep redundancy and test the new capacity. I shard when independent data needs to exceed that practical limit.”

What the answer must demonstrate: Separate physical resources from contention.

Why does doubling application servers not double checkout throughput?

Reveal a model answer

“They may still share the same database, lock, or downstream service. I trace a purchase and measure where time and work accumulate. Adding application capacity helps only the work those instances own; the shared inventory writer may remain the limiting resource.”

What the answer must demonstrate: Find the shared bottleneck.

What does 99.9% availability permit?

Reveal a model answer

“First I would define the measure. Over a 30-day time-based window, 0.1% is 43.2 minutes. Over a million eligible requests, it is 1,000 unsuccessful attempts. These budgets are not interchangeable when traffic changes through the day.”

What the answer must demonstrate: Define eligible and successful requests.

Why include detection time in a recovery plan?

Reveal a model answer

“The customer experiences the outage before the operator starts repairing. If detection takes one minute and verified failover takes six more, checkout is unavailable for seven. I improve both detection and repair and practise the complete sequence.”

What the answer must demonstrate: Measure end-to-end recovery.

Do two copies guarantee durability?

Reveal a model answer

“No. I need to specify when a write is acknowledged, whether the second copy is durable, and which failures it survives. Copies in the same failure domain may disappear together, and a bad deletion can replicate to both. I also need backups and tested recovery.”

What the answer must demonstrate: Name the failure being tolerated.

Is fewer network messages always more efficient?

Reveal a model answer

“No. One message may contain a huge unused payload, while several small messages may run in parallel. I compare bytes, round trips, CPU, and end-to-end latency for the same user operation. Reducing repeated calls can help, but the workload decides.”

What the answer must demonstrate: Count bytes and sequential waits, not just arrows.

What makes a system manageable in an interview answer?

Reveal a model answer

“I show how an operator diagnoses one failed order using a trace identifier and durable states, how alerts reflect failed purchases, and how a rollout can be stopped or reversed. I include schema compatibility and verify recovery rather than ending the design at deployment.”

What the answer must demonstrate: Explain a concrete operator action.

What are a data model, access pattern, invariant, and transaction? How do they guide database choice?

Reveal a model answer

A data model describes the representation: tables, documents, key-value pairs, or graph relationships. An access pattern is a specific query or update, such as recent orders for customer U7. An invariant is a rule that must remain true, such as stock never becoming negative. A transaction treats operations as one logical unit whose changes commit or roll back together. ACID names atomicity, ACID consistency, isolation and durability; the engine and its settings determine the exact guarantees.

For an order service, write down order-by-ID, customer history, and conditional stock allocation. If reducing stock from 5 to 3 must commit with creating a $24 order, a relational database with suitable indexes and a local transaction is a straightforward starting point. Then test expected volume, hot-item contention, and the actual engine's features. SQL and NoSQL labels alone do not determine scale or transaction support.

What the answer must demonstrate: Size alone does not describe a workload.

An order contains its item lines, while product inventory is shared across many orders. Would storing each order as one document make the whole purchase atomic?

Reveal a model answer

“Embedding O81’s lines makes the order read convenient, but MUG9 stock is shared by many orders. Copying available quantity into each order creates competing truths. I would keep stock in one authoritative inventory system and use a supported transaction, or an explicit reservation workflow, to coordinate stock allocation with the order.”

What the answer must demonstrate: Distinguish one aggregate from all shared state.

How does wide-column differ from analytical columnar storage?

Reveal a model answer

“A wide-column model can place U7’s orders in one partition and order them by time for a known serving query. Analytical columnar storage supports scans of selected attributes across many records. Similar names do not make their access shapes or guarantees interchangeable.”

What the answer must demonstrate: Avoid treating column-related names as one category.

A purchase must create order O81 for two $12 items and reduce stock from 5 to 3. Explain ACID for that transaction.

Reveal a model answer

“Atomicity makes stock allocation and order insertion succeed together or have neither change take effect. Correct logic preserves nonnegative stock. Isolation governs concurrent buyers. Durability defines which failures committed O81 survives. I would show the transaction and its settings because saying ‘ACID database’ does not prove the application rule.”

What the answer must demonstrate: Name the rule and distinguish the two meanings.

Two concurrent purchases each request two units when stock is two. What prevents overselling?

Reveal a model answer

“I put UPDATE Inventory SET available = available - 2 WHERE sku = the_requested_sku AND available >= 2 in the same transaction as the order insertion, and require one affected row before continuing. In PostgreSQL Read Committed, the second updater waits and rechecks the predicate. If the first commits stock 2 → 0, the second affects zero rows and rolls back instead of creating an order. A stock CHECK constraint is useful defense, but I still need the transaction and affected-row check.”

What the answer must demonstrate: A fresh read is not an atomic allocation.

Why does a flexible schema still need planning?

Reveal a model answer

“Old and new consumers must agree on quantity, currency, and record versions. Permitting multiple shapes does not tell the application how to interpret them. I would validate required fields and stage compatible readers and writers so a storage change does not silently change meaning.”

What the answer must demonstrate: Flexibility does not eliminate migration work.

Order O81 commits but the response is lost. How should the application recover the outcome?

Reveal a model answer

“The retry carries the same customer-scoped purchase key and request. I claim that unique key when inserting the uncommitted order, before allocating stock. If the key conflicts, I roll back the attempt, then use a fresh transaction to read and validate the original order’s request hash. This returns the original success even if it exhausted the remaining stock. A new purchase ID or a stock check performed before resolving the duplicate would give the wrong retry behavior.”

What the answer must demonstrate: Unknown commit is different from known rollback.

What changes when inventory becomes an independent service?

Reveal a model answer

“The stock and order updates no longer share the original local transaction. I must choose a distributed transaction or durable reservation workflow with explicit intermediate and compensation states. Moving tables across owners without revisiting that boundary loses the guarantee my first design depended on.”

What the answer must demonstrate: Ownership changes can change correctness, not only performance.

What is an index, in plain language?

Reveal a model answer

“It is a maintained search structure that helps locate records without checking every row. An author catalog points to books by an author. In a database, the index stores searchable keys and enough information to find or return matching data.”

What the answer must demonstrate: Explain the read/write tradeoff.

How does an index on (author, title, id) answer author = Le Guin ordered by title?

Reveal a model answer

“It seeks to the first Le Guin entry and scans that contiguous author range in title order. It fetches the matching book rows only if required fields or visibility checks need them, then stops at the range end or limit. The benefit is avoiding unrelated authors, not assuming every query can be served entirely from the index.”

What the answer must demonstrate: Walk the keys rather than naming the structure.

Which index fits customer history sorted newest first?

Reveal a model answer

“I start with customerId, then createdAt descending, then orderId descending for ties. Equality on customer narrows the range and the remaining order supports the requested slice. I would include returned columns only if reducing row lookups justifies a larger index.”

What the answer must demonstrate: Explain equality, ordering, and tie breaking.

When inserting a new book row with ID 15, what additional work do maintained indexes require?

Reveal a model answer

“The table gets a row and each maintained index gets a corresponding entry. The storage engine also performs its logging and any page maintenance required. Extra indexes therefore increase write amplification, memory pressure, and storage even if this insert is only one business operation.”

What the answer must demonstrate: Account for all maintained structures.

What is a covering index?

Reveal a model answer

“It contains the fields needed to answer a query, potentially avoiding separate row fetches. For order history I might include total with the ordering keys. Whether an index-only scan is actually possible also depends on the engine’s visibility rules and query plan.”

What the answer must demonstrate: Do not promise every covered query avoids all table access.

Why can a large OFFSET be expensive?

Reveal a model answer

“The database may still walk past the earlier matching entries before returning the requested page. A keyset cursor lets the next query seek after the last seen ordering tuple. I use a stable tie breaker and define how concurrent inserts affect the browsing session.”

What the answer must demonstrate: Separate ordering and snapshot consistency.

Why might the optimizer ignore an index?

Reveal a model answer

“A query matching 90% of a table may do more work through index-to-row lookups than through a sequential scan; a query matching 100 rows in a million has a different cost. I inspect estimated versus actual rows, buffers, filtering and sort work for the exact query. Small tables, stale statistics and data skew can change the plan.”

What the answer must demonstrate: Avoid absolute rules disconnected from data.

Does an index solve finding data across shards?

Reveal a model answer

“A local index searches within its storage owner. The request still needs to identify the right shard, or query a distributed index or multiple owners. For customer history, customer-based routing and a customer/time local index work together.”

What the answer must demonstrate: Explain routing before local lookup.

Foundation · Question 1Storage engines and data models

What is a storage engine, and how is it different from a data model?

Reveal a model answer

The data model describes records and access semantics, such as messages keyed by room and sequence. The engine organizes their bytes and indexes and performs updates and recovery. B-trees and LSM trees are engine techniques; relational tables and documents are logical models. Choosing SQL does not by itself select a B-tree or define its disk cost.

What the answer must demonstrate: Distinguish the logical interface from physical organization.

Applied · Question 2Storage engines and data models

A B-tree has separators 20 and 50; its middle leaf contains 21, 35, 42, 49. Explain lookup for key 42.

Reveal a model answer

“The root separators guide me to the relevant leaf range, where I find 42’s index entry. Depending on the layout, that entry contains the needed data or points to a separate row. Cached pages can avoid disk reads.”

What the answer must demonstrate: Distinguish logical search steps from physical I/O.

Applied · Question 3Storage engines and data models

An update is acknowledged before its changed data page reaches disk. Under what WAL policy can it survive a process crash?

Reveal a model answer

It can survive when the required recovery records, including the commit decision, were made durable before acknowledgment and recovery correctly replays them. Log-before-data ordering alone does not prove commit-before-ack durability. I must verify the configured synchronization policy and failure model.

What the answer must demonstrate: Name the acknowledgment boundary and failure model.

Foundation · Question 4Storage engines and data models

Why can an LSM contain two values for message 42?

Reveal a model answer

“The old sorted file cannot be changed. An edit first enters a newer memory table and later another file. Reads use the engine’s sequence and snapshot rules to choose the right version. Compaction removes old versions once they are no longer needed.”

What the answer must demonstrate: Explain version visibility, not just file count.

Follow-up · Question 5Storage engines and data models

An LSM contains a tombstone for key 8 and older files may contain key 8’s value. When may the tombstone be removed?

Reveal a model answer

“Only when the engine can prove older values cannot reappear for supported reads and no required snapshot needs that history. Removing the marker merely because it is old can expose an older stored copy.”

What the answer must demonstrate: Logical deletion, compaction and physical erasure differ.

Applied · Question 6Storage engines and data models

What does write amplification of 8 mean at 100 MB/s ingestion?

Reveal a model answer

“With a measurement that includes all the relevant local writes, it implies roughly 800 MB/s of device writes. I would also budget compaction reads, CPU, replication and headroom, and verify the figure under a steady workload.”

What the answer must demonstrate: Define the measurement before multiplying it.

Applied · Question 7Storage engines and data models

Why does an ordered engine not automatically give fast room history?

Reveal a model answer

“The logical key and partitioning still matter. If each full message key is independently hashed to a different shard, a room query fans out. Keeping room and sequence together gives locality but may create a hot room partition.”

What the answer must demonstrate: Connect query shape to both ordering and partitioning.

Follow-up · Question 8Storage engines and data models

How would you test the engine choice?

Reveal a model answer

“I would load representative data, sustain ingestion until compaction reaches normal behavior, and measure tail latency for latest-fifty reads, edits, deletions and recovery. An empty database’s short insert burst hides the deferred maintenance cost.”

What the answer must demonstrate: Evaluate steady-state operation, not only peak foreground throughput.

What does a load balancer do?

Reveal a model answer

“It chooses a healthy backend for incoming service traffic. The client uses one service address; the balancer can route its request to A or B. It distributes work and helps route around detected failures, but shared data and correct write ownership still need their own design.”

What the answer must demonstrate: Explain routing separately from state.

Where do six equal requests go across A, B, and C?

Reveal a model answer

“Under simple round robin: A, B, C, A, B, C. If A has twice the capacity, I can use weights giving A roughly half the traffic. These choices assume requests are similar enough that request count represents work.”

What the answer must demonstrate: Demonstrate a schedule before discussing limitations.

When should a service use L7 routing instead of L4 balancing?

Reveal a model answer

“I need to route by HTTP path, such as /images versus /checkout. An L7 balancer understands those fields, usually after TLS termination. I would also protect the backend connection. An L4 connection balancer is sufficient when I only need transport-level distribution.”

What the answer must demonstrate: Describe what information the layer can inspect.

A has 10 connections, B 2, C 5. Who gets the next one?

Reveal a model answer

“Least connections chooses B if the servers and connection costs are comparable. I would not assume B is least busy if those two connections each contain many expensive streams. I validate the signal against CPU, queueing, and latency.”

What the answer must demonstrate: Qualify the unit of work.

Would sticky sessions solve cart persistence?

Reveal a model answer

“They reduce movement while the chosen server works, but they do not preserve a cart when it fails. I store the authoritative cart durably and use stickiness only if locality improves performance. Then another backend can continue the session.”

What the answer must demonstrate: Separate locality and durability.

A backend B crashes before its next health probe. What happens until the balancer removes it?

Reveal a model answer

“Some requests may still be sent to B until failure detection crosses its threshold. I use bounded timeouts and safe retries. After removal, A and C must have capacity for the redirected work; otherwise detection can turn one crash into a broader overload.”

What the answer must demonstrate: Acknowledge detection delay and correlated failures.

How do you update an instance serving WebSockets?

Reveal a model answer

“I stop new assignments, signal clients to reconnect where the protocol allows, and enforce a drain deadline. Message state lives in durable storage so reconnecting to another instance can resume from a cursor. I cannot assume a balancer transfers the old socket’s process memory.”

What the answer must demonstrate: Explain the long-lived session explicitly.

Does adding a load balancer eliminate all single points of failure?

Reveal a model answer

“No. The balancer, discovery and shared database are separate dependencies. I use redundant balancers with supported traffic failover and enough surviving capacity. If the failed balancer terminated a client connection, that connection may still break: the client reconnects and retries safely. Routing a new request is different from preserving an old connection.”

What the answer must demonstrate: Trace the full failure path.

What is caching? Use product P7 at $20/version 8 to explain the first miss and a subsequent hit.

Reveal a model answer

“Caching keeps a reusable copy to avoid repeating a more expensive operation. Request R1 for product:P7 misses, so the application loads $20/version 8 from the database and stores a copy. The next permitted request R2 hits that copy. The database remains authoritative; the hit is usable only under the page’s freshness and access policy.”

What the answer must demonstrate: Name the source of truth.

When would you choose a local cache rather than a shared one?

Reveal a model answer

“A local memory cache is fast and avoids a network dependency; a local disk cache can hold larger reusable objects. But copies differ across application instances and vanish or become unavailable with the host. A shared cache simplifies sharing at the cost of a network call and another service to operate.”

What the answer must demonstrate: Explain per-instance copies.

Does a 30-second TTL guarantee every read is less than 30 seconds stale?

Reveal a model answer

“Only under specified fill, age, and refresh rules. If a delayed reader fills an already old value with a new 30-second timer, its data age may exceed that bound. I would carry version or source timestamps when the age limit matters and define which moment starts the TTL.”

What the answer must demonstrate: Distinguish cache residency age and data age.

Why can delete-after-write still return the old price?

Reveal a model answer

“Reader R can fetch version 8 before writer W commits version 9, then refill after W deletes the cache. The delete happened, but the late reader resurrected the old copy. I show that timeline and choose either bounded stale display or a stronger version-aware update protocol.”

What the answer must demonstrate: Locate the late refill, then the atomic check.

Why not acknowledge orders from a write-back cache?

Reveal a model answer

“If the cache acknowledges before durable persistence and then loses the entry, the customer can lose an order already reported as saved. I would need a replicated durable log and a tested recovery protocol, or acknowledge only after the required durable commit.”

What the answer must demonstrate: Tie acknowledgment to a loss model.

A two-entry cache receives insert A, insert B, read A, insert C. What do FIFO and LRU evict?

Reveal a model answer

“With two entries and eviction from existing entries, FIFO evicts A because it was inserted first. LRU evicts B because A was accessed more recently. This demonstrates that insertion order and access order are different.”

What the answer must demonstrate: Replay the actual ordering.

Ten thousand readers miss P7 at once. What do you do?

Reveal a model answer

“I allow one refresh for P7 and coalesce the other requests behind it, with bounded waiting. If the product permits it, I serve a stale copy during refresh. I also limit database fallback globally so many different missing keys cannot overwhelm it.”

What the answer must demonstrate: Distinguish same-key and many-key bursts.

How would you add a CDN to an existing image service?

Reveal a model answer

“I keep static objects behind a stable static hostname and point delivery through the CDN. I set origin access, TLS, cache headers, and versioned object paths. On a miss the edge fetches the origin; on a permitted hit it returns its copy. Private objects need a separate authorization-compatible plan.”

What the answer must demonstrate: Explain both migration and key correctness.

What is the difference between a forward and reverse proxy?

Reveal a model answer

“A forward proxy represents clients reaching external destinations, such as employees using a company web gateway. A reverse proxy represents servers to incoming callers, such as shop.example forwarding to an internal order service. Both relay responses back; the names describe role, not one-way packet direction.”

What the answer must demonstrate: Say whose behalf the proxy acts on.

When a reverse proxy terminates HTTPS for an order API, which connection does TLS protect?

Reveal a model answer

“The browser’s TLS connection ends at the reverse proxy, which presents the shop certificate and can inspect the HTTP request. The proxy may then establish a separate protected backend connection. I would not assume browser-to-edge encryption automatically protects the entire path.”

What the answer must demonstrate: Draw both connection segments.

Why can’t the backend trust any X-Forwarded-For value?

Reveal a model answer

“An external caller can send that ordinary header. In a one-edge deployment I replace untrusted claims at the edge with its observed address, and the backend accepts forwarding metadata only from that trusted edge. Multiple proxies require an explicit trusted-hop traversal rule. A forged localhost value must not grant internal access, and an IP address still does not establish user identity.”

What the answer must demonstrate: Separate network provenance and identity.

Can a reverse proxy share an authenticated order response using only its URL as the cache key?

Reveal a model answer

“No. A private order response must not become another user’s response. I choose an authorization-compatible cache policy, often avoiding shared caching for this path. Public versioned product images can use a different policy.”

What the answer must demonstrate: Protect the authorization decision as well as key separation.

Does “open proxy” mean “anonymous proxy”?

Reveal a model answer

“No. Open describes who is allowed to use it; anonymous describes which identifying information it tries to hide. A proxy can be open and still log users or forward identifying headers. Neither term alone establishes privacy or safety.”

What the answer must demonstrate: Treat role, access, and visibility as different dimensions.

The gateway times out on POST /checkout. Can it retry automatically?

Reveal a model answer

“Only if the checkout protocol makes repeating that logical request safe. The origin may already have committed the purchase while the response was delayed. A stable idempotency key and saved result let a retry recover the outcome; an arbitrary new POST may create a second purchase.”

What the answer must demonstrate: A timeout is an unknown outcome.

The origin works but the public subpath fails. What do you inspect?

Reveal a model answer

“I inspect path stripping, relative links, redirects, query strings, and asset routes. If the origin redirects to a root-relative path, it may omit the public prefix. I verify the actual public URL rather than treating origin success as end-to-end proof.”

What the answer must demonstrate: Follow the visible URL through the proxy.

Which responsibilities would you keep out of a generic gateway?

Reveal a model answer

“I can centralize routing, TLS, request-size limits, and some authentication or quota checks. The order service must still check who may read or change an order, and the component committing a purchase must enforce rules such as not selling more stock than is available. Otherwise an alternate internal caller could bypass the only business check.”

What the answer must demonstrate: Explain responsibility and shared failure modes.

Foundation · Question 1Data partitioning and sharding

How is a shard different from a replica?

Reveal a model answer

“A shard owns a different subset of records; a replica is another copy of the same records. A and B split customers, while A1 and A2 could be copies of shard A. I need separate rules for routing to an owner and for keeping that owner’s copies consistent.”

What the answer must demonstrate: Draw ownership and copies separately.

Applied · Question 2Data partitioning and sharding

Why choose customer ID for order partitioning?

Reveal a model answer

“The dominant query asks for one customer’s orders. Keeping those records together allows one routed query and local updates of related order data. I would verify the customer traffic distribution and identify global queries that this choice makes more expensive.”

What the answer must demonstrate: Connect the key to an actual query.

Foundation · Question 3Data partitioning and sharding

When does range partitioning help?

Reveal a model answer

“Queries over adjacent keys can target a small set of contiguous ranges. It is useful when the range matches the query, such as a time slice. The risk is skew: always appending to the newest timestamp range can concentrate writes.”

What the answer must demonstrate: Explain the category and the method.

Applied · Question 4Data partitioning and sharding

Hashing is uniform. Why is one shard still overloaded?

Reveal a model answer

“Uniform placement distributes keys, not necessarily requests. One customer may account for half the work, or one key may be exceptionally large. I inspect traffic and bytes by key, then consider splitting that workload, replicating reads, or allocating dedicated capacity.”

What the answer must demonstrate: Do not promise hashing eliminates hot keys.

Applied · Question 5Data partitioning and sharding

What happens to a join between orders and products?

Reveal a model answer

“If they live on different owners, a local SQL join may no longer cover them. I can perform bounded application lookups, co-locate relevant data, or keep a suitable read copy. For receipts, recording product name and price at purchase time is often the correct historical data.”

What the answer must demonstrate: Distinguish historical facts from current replicas.

Applied · Question 6Data partitioning and sharding

How do you move a shard without losing writes?

Reveal a model answer

“Keep B accepting writes while copying a consistent snapshot tied to log position L0. Apply later logged changes at C. To switch, stop B’s writes and make C apply through B’s final committed position. Then enable C under a new routing version and reject writes using B’s old version. Stale clients refresh their routes and retry the same operation. If I cannot prove B can no longer commit writes, I do not enable C.”

What the answer must demonstrate: Separate data catch-up and ownership transfer.

Applied · Question 7Data partitioning and sharding

What if the shard directory is unavailable?

Reveal a model answer

“Clients can use a cached version only while the ownership protocol makes stale routes safe. Owners validate epochs and reject invalid writes. For metadata changes I need a durable authoritative directory; guessing a new owner can create conflicting histories.”

What the answer must demonstrate: Explain how stale metadata is detected.

Applied · Question 8Data partitioning and sharding

How do you support a report for all orders today?

Reveal a model answer

“Customer-based sharding does not localize a global time query. I can fan out bounded queries and merge results for modest needs, or stream order changes into an analytical store partitioned for reporting. I state the reporting freshness delay and avoid making every checkout wait for analytics.”

What the answer must demonstrate: Name the cost of a query the key does not serve.

Foundation · Question 1Consistent hashing and virtual nodes

What is consistent hashing? Draw a ring and explain why adding a node moves fewer keys than changing a modulo divisor.

Reveal a model answer

Consistent hashing is a placement scheme that limits remapping when owners join or leave. Draw a ring numbered 0–99 with A at 20, B at 50, and C at 80. Hash a key and choose the first clockwise owner, wrapping at 99. Hash 35 belongs to B50; hash 90 wraps to A20.

Add D40: it takes only (20,40] from B, so hash 35 moves to D while hash 45 stays at B. Changing hash(key) mod 3 to mod 4 would change many unrelated assignments. In balanced equal-capacity placement, adding one to N owners moves about 1/(N+1) of keys on average; this particular D40 interval covers 20% of our toy ring. Virtual nodes improve balance, but data still needs migration or cache refill, and one hot key remains a separate problem.

What the answer must demonstrate: Demonstrate the rule with actual positions.

On a 0–99 ring with A20, B50, C80 and keys at 12, 35, 45, 65, 90, which keys move when D40 joins?

Reveal a model answer

“Only P35 moves in our five-key sample. D takes (20,40] from B; P45 is outside that interval and stays with B. A and C keep their existing intervals. I would show the interval, not claim that every key moves to a new server.”

What the answer must demonstrate: Keep a concrete trace distinct from a statistical estimate.

On a clockwise ring with A20, D40, B50, C80, which owner receives B50’s interval when B is removed?

Reveal a model answer

“B’s remaining interval (40,50] passes to C80, the next clockwise owner. P45 moves to C. P35 stays with D. For durable data I must also ensure C obtains the required current state; the placement calculation does not transfer bytes.”

What the answer must demonstrate: Placement and durability are separate responsibilities.

Foundation · Question 4Consistent hashing and virtual nodes

Why not just change hash(key) mod 3 to mod 4?

Reveal a model answer

“That changes many assignments at once, even though most existing machines are still healthy. Hash 35 changes remainder from 2 to 3, while 12 happens to stay at 0. Broad remapping can create expensive migration or cache misses; consistent hashing limits the affected ranges.”

What the answer must demonstrate: Avoid claiming every modulo mapping necessarily changes.

Foundation · Question 5Consistent hashing and virtual nodes

What do virtual nodes improve?

Reveal a model answer

“They give one physical host several separated ring positions, so it owns multiple smaller intervals. With a suitable distribution, this reduces random placement imbalance and can represent differing capacities. It adds token metadata and migration units; it does not create more independent machines.”

What the answer must demonstrate: Count physical failure domains for replication.

One key P35 receives half of all reads. Will more virtual nodes split that hot key?

Reveal a model answer

“No. The same key still maps to one primary owner under this rule. I would consider read replication, caching, or request coalescing, while defining update and freshness behavior. Virtual positions improve distribution across many keys rather than splitting one indivisible key’s traffic.”

What the answer must demonstrate: Key count, bytes, and traffic are different load measures.

How many of 1.2 million keys move when three balanced owners become four?

Reveal a model answer

“The expected share for the new equal-capacity owner is about one quarter, or 300,000 keys. I would label the balance and distribution assumptions. At 500 bytes each that is about 150 MB of payload before overhead, which helps estimate a controlled transfer.”

What the answer must demonstrate: Qualify both arithmetic and assumptions.

A write updates P35 while its ownership moves from B to D. What must the migration protocol guarantee?

Reveal a model answer

“D needs a snapshot and the updates committed while that snapshot is copied. I would catch up, verify, and atomically change the authoritative routing generation under the migration protocol. B must forward or reject stale requests rather than keep an independent writable copy. After D accepts new writes, routing back to B requires reverse catch-up; retaining B’s old snapshot alone does not make rollback safe.”

What the answer must demonstrate: Do not mistake a new ownership map for a complete migration.

Foundation · Question 1Replication and durability

What is replication? How is it different from redundancy and durability?

Reveal a model answer

Replication copies changes to additional replicas. Redundancy is the broader idea of spare resources: a spare machine, disk, or network link can be redundant without containing a usable data copy. Durability is the guarantee that a committed write survives a defined failure set. The replication protocol, durable storage, acknowledgment rule, and failover rules jointly determine that guarantee.

Suppose leader A acknowledges cart v41 before follower B receives it. Replication is configured, but permanently losing A can still lose that acknowledged write. Waiting for the required durable copies reduces this loss exposure while adding network/storage latency and making writes depend on those copies being reachable. Replication also copies a mistaken deletion, so it does not replace a backup.

What the answer must demonstrate: Name the freshness and failure promise.

Foundation · Question 2Replication and durability

Why distinguish received, durable, and applied?

Reveal a model answer

“Received bytes may be only in memory. Durable bytes survive the specified storage failure model. Applied entries are visible to queries. B can have v41 durably logged while ordinary reads still show v40, so acknowledgment and read policy must account for different milestones.”

What the answer must demonstrate: Do not equate a network acknowledgment with query visibility.

Applied · Question 3Replication and durability

A leader acknowledges v41 before a follower receives it, then permanently fails. Explain the possible data loss.

Reveal a model answer

“A responds at .003, fails at .006, and B would receive the change at .008. If A’s storage is lost, the survivors have v40. I either accept that acknowledged-write loss window explicitly or wait for the required durable replica before answering.”

What the answer must demonstrate: Avoid universal durability claims.

Applied · Question 4Replication and durability

With three replicas and a one-replica-loss durability goal, why might the commit protocol wait for two durable copies instead of all three?

Reveal a model answer

“Two durable copies leave at least one copy of an acknowledged entry after any one participant is lost. With a safe election and commit protocol, the surviving majority preserves that committed history and can continue. Waiting for all three adds a copy but makes the slowest replica control acknowledgment and stops writes if any replica is unreachable. I would choose two only because it meets the stated one-failure contract; the count alone is not the safety proof.”

What the answer must demonstrate: Failure budget and acknowledgment must agree.

Applied · Question 5Replication and durability

A write of v41 succeeds, but a subsequent session read returns v40. What should you inspect?

Reveal a model answer

“Check which replica answered and how far it had applied the write log. It may have saved v41 without making it readable yet. To read my own write, use the verified current leader or wait for a follower to apply the returned commit position. That position must still identify the right history after failover. A former leader or an arbitrary application version cannot prove freshness.”

What the answer must demonstrate: Waiting a fixed time does not prove that the required update is visible.

Foundation · Question 6Replication and durability

How do replicas and shards fit together?

Reveal a model answer

“A shard owns a subset of records, while replicas store copies of that subset. Cart C17 can belong to one shard with three replicas. Adding shards can divide data and write work; adding followers preserves copies and can spread eligible reads. Each follower still has to process its shard’s write stream.”

What the answer must demonstrate: Do not count duplicated processing as partitioned work.

Follow-up · Question 7Replication and durability

A new leader B takes over from isolated leader A. What prevents A from continuing to commit writes?

Reveal a model answer

“A must lose the ability to commit new writes when B takes over. Missing heartbeats alone does not prove A stopped. Use the database’s safe election and fencing protocol to reject the old leader, then update routing so clients find B.”

What the answer must demonstrate: Routing is discovery, not ownership enforcement.

Follow-up · Question 8Replication and durability

Every replica contains a mistaken deletion. What next?

Reveal a model answer

“I stop the faulty job, restore retained history in isolation, identify C17’s last valid state, and verify the repair. Promoting another current replica cannot undo a deletion they all copied correctly. I would also check the full affected range.”

What the answer must demonstrate: Replicas and recovery history solve different failures.

What is the CAP theorem? Define C, A, and P, and explain the triangle with a concrete example.

Reveal a model answer

CAP says a distributed read/write system cannot guarantee both linearizable consistency and completion of every request to a nonfailed participant when network partitions are allowed. C means clients observe one up-to-date copy: after a write completes, a later read must return it or a newer write. Formally, operations fit one valid history respecting real-time order. A means every such request eventually completes according to its contract. P means live replicas can be unable to exchange messages.

Draw C, A, and P at the triangle's vertices. Label CP as preserving one history while some operations wait or fail, AP as permitting completion with weaker consistency, and CA as requiring that partitions are excluded from the guarantee. Do not present P as a network failure you can disable in production.

For example, East and West both store S7 as free. They lose contact. East confirms client A's reservation. A later West read cannot learn that fact: returning free violates C; refusing or waiting without completion gives up A. The design should state which behavior is acceptable for that operation.

What the answer must demonstrate: State the theorem before the caveats; define all three letters and use one completed-write/later-read partition trace.

Why does a quick 503 not prove CAP availability?

Reveal a model answer

“The client reached a working participant but did not complete the requested seat read. The server replied quickly, which is useful operationally, but refused the object operation. I would count that separately from a valid ‘already reserved’ result and separately from the product’s latency target.”

What the answer must demonstrate: Separate infrastructure failure from legitimate business rejection.

Can a partition happen while both databases are healthy?

Reveal a model answer

“Yes. East and West may both run normally and answer their local clients while network messages between them are dropped. That is why checking each process’s health is insufficient. I need to know which communication and authority assumptions an operation requires.”

What the answer must demonstrate: A network partition is not necessarily a server crash.

East and West start with S7 free, then become partitioned. East confirms a reservation at 10:00:02; a West read begins at 10:00:03. Why can West not guarantee a linearizable answer while completing every such read?

Reveal a model answer

“West has the same local state in several possible histories: client A reserved in East, someone else reserved, or nobody wrote. No East message has arrived. Its old null value cannot distinguish them. Answering immediately may choose the wrong history; waiting for information can prevent completion during a continuing partition.”

What the answer must demonstrate: Explain the missing information, not just repeat ‘choose two.’

How would you handle the last seat during a partition?

Reveal a model answer

“I would allow only the participant with valid write authority to perform the atomic available-to-reserved transition. A disconnected minority would decline it. That may stop some purchases, but a successful confirmation then means the seat was reserved by the node currently authorized to make that decision. I would specify the quorum and safe leader change rather than relying on a product label.”

What the answer must demonstrate: Adding replicas is not the same as defining a safe election protocol.

Does preventing double sales imply every read is CAP-consistent?

Reveal a model answer

“No. I can send all reservations through one atomic authority while serving a stale seating map elsewhere. The business invariant can hold even when that display is not linearizable. Conversely, a correctly ordered store can still oversell if my application uses an unsafe read-then-write algorithm.”

What the answer must demonstrate: CAP C and application invariants are related design concerns, not identical definitions.

A reservation request times out without a known outcome. How should the client retry?

Reveal a model answer

“Reuse the operation identifier and ask the authority for the durable outcome. A timeout means the response was not received; it does not prove the reservation failed. If the old attempt committed, return that result. If it did not, process the retry under the same ownership rules.”

What the answer must demonstrate: A missing response is an unknown outcome.

What must happen after the partition heals?

Reveal a model answer

“Replicas must converge on the protocol’s authoritative history, and obsolete writers must remain fenced. I would verify catch-up before routing reads that promise current state. If our policy allowed conflicting writes, I also need an explicit business repair policy; network recovery alone cannot choose who deserves a promised seat.”

What the answer must demonstrate: Recovery must honor promises made before and during the fault.

Foundation · Question 1Consistency models

What is a consistency model? Explain it using a write of version 11 followed by a read.

Reveal a model answer

A consistency model defines the read results and operation orders a system allows. If client A completes a write of version 11 and client B then reads, linearizability forbids the old version 10 when no other write intervened. Eventual consistency may temporarily allow version 10. The choice describes a visible contract, not whether the title text is factually correct.

What the answer must demonstrate: Define permitted observations and the object or transaction scope.

Foundation · Question 2Consistency models

An interviewer says “the system must be consistent.” Which meaning should you clarify?

Reveal a model answer

I ask whether the requirement concerns read visibility or a business invariant. For CAP consistency, a write that completes before a read starts must be visible to that read, or superseded by a newer write. More generally, I name the required consistency model, such as linearizable or causal. ACID consistency means transactions preserve rules such as nonnegative stock. I would state the operation and show a concrete forbidden result.

What the answer must demonstrate: Connect the familiar current-value explanation to the formal model, and keep ACID validity separate.

Applied · Question 3Consistency models

A write from v10 to v11 overlaps a read on another client. Must a linearizable read return v11?

Reveal a model answer

“Not necessarily. Under linearizability the read may take effect before or after the concurrent write. I would inspect invocation and response intervals; a read beginning after the write completed is the clearer test.”

What the answer must demonstrate: Do not replace the definition with a vague latest-value rule.

Foundation · Question 4Consistency models

Give a history allowed by sequential consistency but not linearizability.

Reveal a model answer

“Client A completes writing v11, then an independent client B starts a read and gets v10. With no other operations, a total order can put client B’s read first, preserving each client’s order. Real-time completion forbids that placement under linearizability.”

What the answer must demonstrate: Keep process order separate from wall-clock order.

Applied · Question 5Consistency models

How do you stop replies appearing before their comments?

Reveal a model answer

“I attach the parent or a sufficient dependency context to client B’s reply. A replica cannot expose the reply until it can expose that history. That is a visibility rule, not just sorting by arrival timestamp.”

What the answer must demonstrate: Dependencies do not imply a total order for independent writes.

Applied · Question 6Consistency models

How can a session preserve read-your-writes when failing over from a replica at v11 to one at v10?

Reveal a model answer

“The save response carries a storage position or version context. The next server must prove it has applied that context before answering. If it cannot, it routes or waits; silently returning v10 violates the session promise.”

What the answer must demonstrate: Describe failover as well as the normal request path.

Follow-up · Question 7Consistency models

Does a five-second TTL guarantee data no older than five seconds?

Reveal a model answer

“Only under additional assumptions. If a cache fills from a replica already thirty seconds behind, a fresh cache entry is still stale. I need an authoritative reference, propagation limits, and behavior when the bound cannot be met.”

What the answer must demonstrate: Cache age and source age differ.

Follow-up · Question 8Consistency models

Does causal consistency merge concurrent title edits?

Reveal a model answer

“No. It tells us which edits depend on which earlier edits. Independent edits still need a conflict policy, such as preserving both versions for the user or a domain-specific merge. A last-writer rule chooses a winner but can lose intent.”

What the answer must demonstrate: Separate causal ordering, convergence, and application semantics.

Applied · Question 9Consistency models

Why not require linearizability for every read in a collaborative application?

Reveal a model answer

“It may be acceptable, especially at modest scale, but I would compare the added coordination latency and the operations that may become unavailable with what the product requires. The editing session and comment dependencies can often have clear weaker contracts, while ownership still needs stricter enforcement.”

What the answer must demonstrate: Make the choice per operation, not per marketing category.

Foundation · Question 1Transaction isolation

What is transaction isolation, and why is it different from atomicity?

Reveal a model answer

Atomicity makes a transaction’s changes commit together or abort together. Isolation controls how concurrent transactions observe and interfere with one another. In the roster example, two atomic leave requests can both read 2 and update different rows, leaving 0 on duty under snapshot isolation. The database needs a concurrency rule that protects the shared business condition.

What the answer must demonstrate: Separate all-or-nothing changes from safe concurrent decisions.

Foundation · Question 2Transaction isolation

Explain nonrepeatable and phantom reads without jargon.

Reveal a model answer

“A nonrepeatable read occurs when transaction T1 reads row D2 twice and observes different committed values because another transaction changed it. A phantom occurs when T1 repeats a predicate query, such as all on-duty rows, and sees a new or missing matching row. The first concerns an existing row’s value; the second concerns membership in a result set.”

What the answer must demonstrate: Use one row versus a matching set.

Applied · Question 3Transaction isolation

Does Repeatable Read prevent phantoms?

Reveal a model answer

“I would ask which database. The SQL standard’s minimum guarantees allow that phenomenon, while PostgreSQL Repeatable Read uses a stable snapshot and prevents it. Neither statement means PostgreSQL Repeatable Read prevents our write-skew example.”

What the answer must demonstrate: Do not generalize product behavior from the level name.

Applied · Question 4Transaction isolation

Two transactions read revision 8 and both assign 9. How do you prevent this lost increment?

Reveal a model answer

“I use an atomic increment or a compare-and-update against the expected revision, checking whether it succeeded. Reading 8 in application code and later assigning 9 in both requests loses one increment.”

What the answer must demonstrate: A local race fix must cover the business decision to enforce it.

Applied · Question 5Transaction isolation

To protect count(on_duty) >= 1 using a shared guard row, when must the guard be locked relative to reading the count?

Reveal a model answer

“Before reading the state used to decide whether someone may leave. I use a transaction pattern whose post-lock query observes the previous holder’s committed result; with Read Committed, a subsequent query gets a fresh statement snapshot.”

What the answer must demonstrate: Lock timing and snapshot timing must agree.

Follow-up · Question 6Transaction isolation

T2 receives a serialization failure. What does the application do?

Reveal a model answer

“Abort the failed attempt and retry the complete transaction: reads, validation, and writes. If another transaction reduced the on-duty count to one, the new execution must reject the off-duty transition. Retrying only the final write reuses an invalid decision.”

What the answer must demonstrate: Retries must recompute the decision.

Follow-up · Question 7Transaction isolation

How do you emit an off-duty notification only for a committed transition when its transaction may abort and retry?

Reveal a model answer

“I record the notification intent atomically with the successful roster transaction. A separate worker sends it using a stable event identifier. The retried transaction body must not perform irreversible external work.”

What the answer must demonstrate: Explain which database changes commit together and which later message delivery still needs deduplication.

Applied · Question 8Transaction isolation

What operational costs should you measure for a guard row that serializes all changes to one roster?

Reveal a model answer

“I measure wait time, transaction length and contention by roster. A long-held guard is a latency bottleneck even if CPU looks idle. I keep the protected work short and test simultaneous leave, deletion and transfer operations.”

What the answer must demonstrate: More concurrency can worsen a serialized bottleneck.

Define quorum, consensus, lease, and fencing. Which problem does each solve?

Reveal a model answer

A quorum is a required response set, such as two of three controllers. Consensus makes those controllers agree on an ownership decision or committed log despite the failures it tolerates. A lease gives ownership for a limited interval. Fencing adds an increasing ownership token that the output store checks atomically with a write.

For export E9, controllers grant W1 epoch 7. W1 pauses; the lease expires; controllers agree to grant W2 epoch 8. W2 publishes with token 8. If W1 resumes and presents 7, the store rejects it. The lease did not stop W1's CPU from executing; the fencing check stops its stale effect after newer authority reaches the resource. Quorum overlap helps the agreement proof but does not supply a complete consensus protocol.

What the answer must demonstrate: Start from actual sets rather than a memorized equation.

Why does overlap not prove linearizability?

Reveal a model answer

“With three replicas, read groups {R1,R2} and {R2,R3} do overlap at R2. But suppose only R1 saw an incomplete write of v9 while R2 and R3 still have v8. A first read returns v9 from R1, then a later read through R2 and R3 returns v8 without another write. The problem is that the read exposed a value without preserving it for later reads. Quorum intersection alone does not define safe version selection, write-back, commitment, or recovery.”

What the answer must demonstrate: Use overlapping replica sets and non-overlapping-in-time reads; explain why a selected value must remain visible to later reads.

What does consensus provide when three controllers assign one owner for export job E9?

Reveal a model answer

“It gives the controllers one agreed sequence of ownership transitions, so W1 expiry and W2’s epoch-8 grant are not independently invented on different copies. I would use a proven replicated-log protocol whose election and commit rules preserve the history after controller failure.”

What the answer must demonstrate: Agreement on metadata does not atomically include every external effect.

What happens when two of three controllers are unreachable?

Reveal a model answer

“Only one remains, so the majority protocol cannot safely advance ownership. I would stop new grants and report reduced availability. I would not let the isolated replica infer that its stale state is now authoritative because it is the only one this client can reach.”

What the answer must demonstrate: Distinguish safety from continued progress.

W1 has fencing token 7; replacement W2 publishes with token 8. W1 resumes. What must the output store check?

Reveal a model answer

“W2’s publication has fence 8, so the output store has atomically recorded that generation with the manifest. W1 arrives carrying 7. The store rejects 7 before changing the protected state, preventing W1 from replacing W2’s newer result.”

What the answer must demonstrate: A separate preflight check leaves a race.

Does a fence instantly revoke old work everywhere?

Reveal a model answer

“Not necessarily. A resource comparing against its latest accepted fence learns about generation 8 when that newer authority reaches it. It prevents older writes after that point. If the requirement is immediate revocation everywhere, I need current-ownership validation or another stronger coordinated boundary.”

What the answer must demonstrate: Describe the precise fencing guarantee rather than implying physical process termination.

A valid worker retries publication with the same fencing epoch 8. Why is an operation idempotency key still needed?

Reveal a model answer

“Epoch 8 says W2 is an eligible owner. It does not distinguish one publication attempt from a retransmission of the same attempt. I use a stable publication ID so a lost response does not create duplicate effects while that ownership is still valid.”

What the answer must demonstrate: Operation identity and authorization are independent checks.

What do you reconcile after the controller outage ends?

Reveal a model answer

“I inspect the committed ownership history, the output store’s accepted fence and manifest, and any unfinished files. A worker saying it finished is weaker than the protected publication record. I then resume or retry with stable IDs and valid ownership rather than blindly rerunning every reported job.”

What the answer must demonstrate: Recovery must consult the state that actually governs the external result.

Foundation · Question 1Idempotency, retries, and timeouts

What does idempotency mean, and why does a timeout make it useful?

Reveal a model answer

Idempotency means repeating the same logical operation has the same intended effect as doing it once. A timeout does not prove failure: O17 may have committed before its response was lost. U9 retries buy-204 with the same caller and payload so the service returns O17 instead of making another purchase.

What the answer must demonstrate: Distinguish one business effect from one transport attempt.

Foundation · Question 2Idempotency, retries, and timeouts

What exactly does the idempotency key identify?

Reveal a model answer

“One caller’s logical intent, such as U9’s purchase attempt buy-204. I scope it to the authenticated caller and compare a canonical payload fingerprint. A fresh network retry reuses it; a new intended purchase gets a different key.”

What the answer must demonstrate: Separate caller, intent, and payload.

Two requests both see no saved result. How is one order guaranteed?

Reveal a model answer

“The claim and business effect must share an atomic boundary, such as a unique request record and order insert in one transaction. A separate check-then-insert can let both proceed. The losing concurrent attempt waits for or retrieves the winner’s outcome.”

What the answer must demonstrate: Show the atomic boundary.

Can idempotency records expire after a minute?

Reveal a model answer

“Only if the contract prevents valid retries after that minute or another durable identity prevents repetition. Deleting the record can make a delayed duplicate look like a new operation. I align retention with the retry horizon, business identifiers, and downstream retention.”

What the answer must demonstrate: Treat deduplication retention as part of correctness.

Why doesn’t a local transaction make the external charge exactly once?

Reveal a model answer

“The provider is outside that transaction. It can charge and lose its reply before we save the result. I use one durable provider attempt key, query or reconcile its outcome, and process duplicate notifications safely. The local order and remote charge have separate commit boundaries.”

What the answer must demonstrate: Avoid blanket exactly-once claims.

Three layers each make three attempts. What reaches the bottom?

Reveal a model answer

“In the worst simple nesting, up to 27 calls for one user operation. That extra work can keep a struggling dependency down. I choose one retry layer or a shared budget, cap attempts and total time, and stop retrying permanent failures.”

What the answer must demonstrate: Show the multiplication and the bound.

How do timeouts relate to a two-second user budget?

Reveal a model answer

“The two-second budget becomes an absolute deadline two seconds after the request starts. Each downstream call gets at most the remaining time as its timeout, including planned retries. Independent two-second waits at every layer can greatly exceed it and keep doing work after the user has left.”

What the answer must demonstrate: Distinguish stopping work from reversing it.

How do you stop a slow export dependency from taking down checkout?

Reveal a model answer

“I separate concurrency pools so exports cannot consume all checkout workers or connections. I bound queues and use deadlines, then reject or defer lower-priority work when capacity is exhausted. A breaker can limit calls to the unhealthy dependency while probing recovery.”

What the answer must demonstrate: Protect a finite resource and explain overload behavior.

What is a message queue? Define producer, consumer, asynchronous completion, and backpressure.

Reveal a model answer

A producer submits a message; a broker stores and delivers it; a consumer processes it. A message queue buffers that work so its execution can happen asynchronously, after the initial request has durably accepted it. Accepted and completed are different states. A photo upload can return a job ID while its thumbnail is still pending.

Backpressure limits incoming or concurrent work when downstream processing cannot keep up. If arrival is 600 jobs/s and workers can process 400 jobs/s, the backlog grows by 200 jobs/s. Bound admission or increase processing capacity instead of treating an unbounded queue as a solution. For reliable handoff, an outbox can commit the photo and job intention together; consumers still need duplicate-safe effects because delivery or acknowledgment can repeat.

What the answer must demonstrate: Acceptance and completion are different promises.

How does a work queue differ from a retained event log?

Reveal a model answer

“The work queue assigns J501 and manages retry eligibility. A retained log lets consumers replay an ordered sequence from their positions, possibly in independent groups. I would choose based on work assignment versus replay needs and inspect the real retention and delivery semantics.”

What the answer must demonstrate: Do not infer guarantees from product-category names.

A database commits photo P501, then the process dies before publishing its job J501. How does a transactional outbox close that gap?

Reveal a model answer

“I store the photo and an outbox intention in one transaction. A relay can find committed unpublished intentions after recovery. This prevents P501 being accepted without a durable record that processing must happen, while avoiding a false claim of one transaction across database and broker.”

What the answer must demonstrate: Outbox solves a missing handoff, not every duplicate.

The worker commits the result and crashes before acknowledging. Walk the retry.

Reveal a model answer

“The job is delivered again after the broker did not record completion. A worker reads the durable job receipt and returns the established outcome. If two duplicate workers race before either receipt exists, a unique job-ID constraint and one result transaction choose the winner; the loser rolls back and reads that winner’s outcome. A preflight receipt lookup alone would not prevent two concurrent effects.”

What the answer must demonstrate: Deduplication placement determines correctness.

Does a consumer’s local deduplication receipt make an external API call exactly once?

Reveal a model answer

“No. A remote object write is outside the result database transaction. I give each render attempt an immutable object key and use the stable photo/version/recipe identity for the logical job. The database transaction chooses one reference and records its receipt; it never lets a duplicate overwrite the chosen bytes. Cleanup cannot delete an active attempt that may still publish. A different provider effect needs that provider’s idempotency contract or reconciliation of unknown outcomes.”

What the answer must demonstrate: Local atomicity does not automatically include a remote effect.

A thumbnail job for photo version 2 finishes after version 3 is published. What should the result commit check?

Reveal a model answer

“The authoritative ready update checks which photo version it belongs to. A stale version-2 result may be retained or cleaned up, but it cannot replace the version-3 reference. Ordering events by photo can help, yet the conditional update protects against retry and completion reordering.”

What the answer must demonstrate: Stable identity must represent stable intent.

For 60 seconds, arrivals are 600 jobs/s and workers complete 400/s. Arrivals then fall to 200/s. Calculate backlog growth and drain time.

Reveal a model answer

“Six hundred arrivals minus four hundred completions gives two hundred extra jobs per second. Over sixty seconds that is twelve thousand jobs. When arrivals fall to two hundred, the spare capacity is two hundred, so recovery takes about sixty seconds under stable-rate assumptions.”

What the answer must demonstrate: Use service rate minus arrival rate when estimating drain time.

J501 contains an image that can never be decoded. Should it retry forever?

Reveal a model answer

“No. I would classify the permanent failure, stop after a bounded policy, store diagnostics, and expose a failed state or dead-letter workflow. Infinite retries consume capacity and can delay valid work. Any operator replay should be controlled and preserve job identity and version rules.”

What the answer must demonstrate: Bound both retry effort and user-visible delay.

Foundation · Question 1Distributed transactions and sagas

What problem does a distributed transaction or saga solve?

Reveal a model answer

A distributed transaction spans multiple transactional participants and needs a coordinated commit-or-abort outcome. A saga addresses a related business need through separately committed local transactions and compensation. For O81, creating an order, reserving 2 mugs, and authorizing $24 can succeed or fail separately. A capable 2PC system coordinates one commit decision; a saga records local progress and compensates failures. I first ask whether the work could remain in one simpler database transaction.

What the answer must demonstrate: Identify the actual independent commit boundaries.

Foundation · Question 2Distributed transactions and sagas

What does a yes vote in 2PC mean?

Reveal a model answer

“The participant has prepared enough durable state and retained the necessary protections to honor a later commit decision. It is stronger than saying the request looks valid right now.”

What the answer must demonstrate: Prepared is a durable protocol state, not a best-effort check.

Follow-up · Question 3Distributed transactions and sagas

Does 2PC guarantee serializable transactions?

Reveal a model answer

“2PC coordinates the final commit or abort outcome. Isolation depends on the concurrency-control protocol over the affected reads and writes. I would not claim serializability just because every participant votes on one decision.”

What the answer must demonstrate: Atomic commit and isolation solve different parts of correctness.

A payment authorization A81 times out with no known result. What should a durable workflow do next?

Reveal a model answer

“Save the outcome as unknown and use A81 to check with the provider. I do not create A82 just to retry: A81 may already have succeeded.”

What the answer must demonstrate: Do not promise exactly-once effects across an unsupported boundary.

An inventory hold expires at 120 seconds and payment authorization succeeds at 125. Can the order be confirmed?

Reveal a model answer

“Not from the authorization alone. Inventory must atomically verify or convert a valid hold, and H81 is expired. I keep confirmation conditional and void the authorization while cancelling the order.”

What the answer must demonstrate: Two authorities must enforce their own conditions.

What happens if the compensating void also fails?

Reveal a model answer

“The cancellation has an outstanding cleanup state with a stable void identifier. A worker retries or queries it, and an age-based alert exposes work that cannot finish automatically.”

What the answer must demonstrate: Do not hide unfinished compensation behind a terminal label.

Follow-up · Question 7Distributed transactions and sagas

What guarantee does a transactional outbox add to a distributed workflow?

Reveal a model answer

“It atomically records the local state transition and the intent to send the next message. After a crash, the relay can find that intent. The relay may publish twice, so consumers still need idempotent handling.”

What the answer must demonstrate: Keep the outbox guarantee within its actual transaction boundary.

Would you use orchestration or choreography for an order workflow with inventory holds, payment authorization, and compensation?

Reveal a model answer

“I would start with an explicit coordinator because the order’s deadlines, compensation and user-visible status form one workflow that operators must inspect. Services still own inventory and authorization details.”

What the answer must demonstrate: Explain operational ownership instead of declaring one style universally better.

Compare short polling, long polling, SSE, and WebSocket. What does real-time mean for a chat application?

Reveal a model answer

For chat, define a freshness target such as new messages normally appearing within 300 ms; this is an illustrative product target, not a property automatically guaranteed by a transport. Short polling repeats a request on a timer, creating idle traffic and up to roughly one interval of waiting. Long polling holds a request until data arrives or it times out, then the client starts another. SSE keeps an HTTP response open for text events from server to client. WebSocket maintains a full-duplex framed message channel.

For infrequent notifications, polling may be sufficient. For mostly one-way live updates, SSE plus ordinary HTTP commands can be simple. For frequent chat messages, typing, and acknowledgments in both directions, WebSocket is a reasonable choice. All choices need authentication, bounded buffering, reconnect, and a durable cursor/history policy; the socket alone cannot restore missed messages.

What the answer must demonstrate: Separate application exchange lifecycle from underlying connection reuse.

100,000 clients poll every five seconds. An event arrives at :02 between polls at :00 and :05. Estimate idle QPS and event delay.

Reveal a model answer

“100,000 clients divided by a five-second interval produce 20,000 requests/s even without updates. An event at :02 waits three seconds until the :05 poll, plus network and processing time. Uniformly timed arrivals wait roughly half an interval on average.”

What the answer must demonstrate: State the arrival and interval assumptions.

What exactly is held during long polling?

Reveal a model answer

“The server holds one HTTP request until an update exists or the timeout expires. For example, a request after cursor 500 returns event 501, and the client immediately requests after its applied cursor again. The response may contain a batch; long polling means one response per request, not necessarily one event. History bridges the short gap before the next held request.”

What the answer must demonstrate: Explain wait, response, reissue, and timeout.

Describe how the WebSocket channel begins.

Reveal a model answer

“For HTTP/1.1, the client requests an upgrade and the server validates it and returns 101 before exchanging WebSocket frames. HTTP/2 and HTTP/3 have extended-CONNECT mechanisms when supported. I would specify what our gateway and clients actually support, authenticate the session, validate browser Origin, and authorize subscriptions; protocol negotiation alone grants no user permission.”

What the answer must demonstrate: Handshake, authentication, and durability are distinct mechanisms.

Could the receiving client send messages while receiving SSE?

Reveal a model answer

“Yes. The receiving client can receive a continuing event-stream response and send commands through separate HTTP POST requests. SSE is one-way on that stream, not a prohibition on the browser making other requests. It is attractive when live traffic is primarily server-to-client.”

What the answer must demonstrate: One-way stream does not mean one-way application.

A client receives event 501 but reconnects with last-applied cursor 500. How should replay work?

Reveal a model answer

“Replay event 501 from durable history and apply it idempotently by message ID. Cursor 500 must mean the application applied every event through that position in the relevant stream. Native SSE Last-Event-ID can advance before the handler durably applies an event, so I would use the explicit application cursor for this stronger replay contract. A gateway writing bytes is not evidence that the recipient recorded the update.”

What the answer must demonstrate: Connection delivery and application progress can differ.

How should a gateway handle a receiving client that consumes events slower than they arrive?

Reveal a model answer

“I bound the outgoing buffer. Depending on the event contract, I can drop optional typing updates, summarize state, or disconnect and resume durable messages later. I cannot let one slow client grow gateway memory without limit.”

What the answer must demonstrate: Separate replaceable hints from durable events.

Why can a simultaneous reconnect after a gateway outage cause another outage?

Reveal a model answer

“A large connection outage can cause every client to reconnect and replay simultaneously. I would use randomized retry delays, admission control, and bounded replay work while protecting the history store. A healthy gateway fleet can still overload its shared dependencies during recovery.”

What the answer must demonstrate: Recovery traffic and permission changes are part of the protocol.

Foundation · Question 1Probabilistic data structures

What is a probabilistic data structure? Use a Bloom filter to explain its possible error.

Reveal a model answer

It uses randomization to obtain a useful space or performance tradeoff. Some probabilistic structures answer exactly; the Bloom filter is an approximate membership summary with a defined error model. In the Bloom example, A sets bits 2 and 7 and B sets 7 and 12. C tests 2 and 12, so the filter says possibly present even though C was never inserted: a false positive. It no longer knows which item set each bit.

What the answer must demonstrate: Name the supported question, error direction, and business consequence.

Applied · Question 2Probabilistic data structures

Under what coverage and update assumptions is a Bloom-filter negative safe to trust?

Reveal a model answer

“It proves absence from a correctly maintained filter’s inserted set. To infer absence from the database, the filter must cover that database state. A stale or interrupted rebuild may omit real entries.”

What the answer must demonstrate: State which set the guarantee describes.

Applied · Question 3Probabilistic data structures

Estimate memory for one million URLs at 1% false positives.

Reveal a model answer

“Using the standard idealized formulas, I need about 9.59 million bits, or 1.20 decimal MB, and about seven hash positions per item. I would add implementation overhead and headroom for growth.”

What the answer must demonstrate: Keep bits and bytes distinct and acknowledge the sizing assumptions.

Applied · Question 4Probabilistic data structures

Of 100,000 membership checks, 80% are absent. With a 1% Bloom false-positive rate, how many exact preliminary reads remain?

Reveal a model answer

“Of 100,000 checks, 80,000 are absent. At a 1% false-positive rate about 800 absent checks still reach the database, alongside 20,000 present checks. That is about 20,800 reads instead of 100,000.”

What the answer must demonstrate: Do not confuse lookup reduction with eliminating all authoritative work.

Follow-up · Question 5Probabilistic data structures

Bloom key A sets bits 2 and 7; B sets 7 and 12. Why can deleting A not simply clear its bits?

Reveal a model answer

“B shares bit 7, so clearing it can turn B into a false negative. Ordinary Bloom bits do not record ownership. I need a correctly managed counting variant or a rebuild/epoch policy.”

What the answer must demonstrate: Deletion changes the guarantee unless ownership is accounted for.

Foundation · Question 6Probabilistic data structures

Could HyperLogLog replace the visited-URL set?

Reveal a model answer

“No. HyperLogLog estimates distinct cardinality; it cannot answer whether a particular URL was seen or enumerate URLs. It is useful for aggregate crawler statistics, while exact claim decisions need an exact set or database.”

What the answer must demonstrate: Separate an aggregate estimator from a membership structure.

Follow-up · Question 7Probabilistic data structures

Why does Count-Min take the smallest counter?

Reveal a model answer

“Each counter contains the item’s own increments plus collisions. Under nonnegative insert-only updates, taking the minimum reduces collision inflation without dropping below the true count. In the example, min(27,23,22) estimates a true count of twenty as twenty-two.”

What the answer must demonstrate: Qualify the update model and distinguish estimation from enumeration.

Applied · Question 8Probabilistic data structures

Can two crawler workers merge their sketches?

Reveal a model answer

“Yes, when the sketch types, dimensions, hash functions and item normalization are compatible. Bloom union uses OR; HyperLogLog uses register maxima; Count-Min sums counters.”

What the answer must demonstrate: Compatible arrays are not enough; semantics must match.

Foundation · Question 1Keyword search and vector retrieval

What is an inverted index? If reset maps to {D2,D4} and access to {D1,D4}, how does reset AND access execute?

Reveal a model answer

An inverted index maps a term to the documents containing it. Here reset maps to D2 and D4, while access maps to D1 and D4. Intersecting the posting lists returns D4 without scanning every document body. Positions support phrases and term statistics support ranking.

What the answer must demonstrate: Build the two lists and distinguish lexical matching from similarity.

When does vector search help?

Reveal a model answer

“It helps retrieve semantically related wording, such as lost phone matching authenticator recovery. It is weaker for some precise identifiers and does not establish truth or permission, so I evaluate it alongside lexical search and metadata filters.”

What the answer must demonstrate: Similarity is a retrieval signal.

What does approximate nearest-neighbor search trade away?

Reveal a model answer

Approximate search can miss neighbors that an exact result would include, in exchange for less work on suitable workloads. I compare it with an exact baseline under the same metric and eligibility filters, then tune latency, memory and recall together. Exactness does not require a full scan if an index can safely prove which candidates cannot win.

What the answer must demonstrate: Separate approximation quality from semantic quality.

Estimate raw storage for ten million 768-dimensional float32 vectors.

Reveal a model answer

“Each vector is 768 × 4 = 3,072 bytes. Ten million require 30.72 GB in decimal units before graph links, metadata, text, and replicas. I would size those separately and benchmark any quantization loss.”

What the answer must demonstrate: Keep units and overhead explicit.

Why not add a keyword score directly to a cosine score?

Reveal a model answer

“Their scales and distributions differ. I can calibrate a learned combination or start with rank fusion, then evaluate. Reciprocal rank fusion uses positions in each result list and avoids pretending unlike raw scores have the same meaning.”

What the answer must demonstrate: Candidate recall bounds reranking.

Why can filtering the final top twenty return no useful result?

Reveal a model answer

“All twenty may belong to another tenant even though relevant authorized documents exist deeper in the collection. I apply an eligible-document retrieval strategy and evaluate selective filters. In every case I enforce authorization before content leaves the trusted retrieval boundary.”

What the answer must demonstrate: Distinguish candidate starvation from data exposure.

Three of five returned documents are relevant; ten relevant documents exist. What are precision and recall?

Reveal a model answer

“Precision@5 is 3/5, or 60%. Recall@5 is 3/10, or 30%. I also measure ranking quality because users often inspect only the first results.”

What the answer must demonstrate: Use the correct denominator.

A document was deleted but remains searchable. How do you fix the contract?

Reveal a model answer

Send versioned deletion markers to the index and measure cleanup delay. To block access immediately, do not rely only on that delayed index update. Check current source permissions for the exact document and policy version before fetching its body or sending it to a model. Fetch that fixed content version; if versions differ, check permission again. Apply the promised access check when releasing the response too.

What the answer must demonstrate: Treat freshness and authorization as explicit guarantees.

What is the difference between authentication, authorization, and tenant isolation?

Reveal a model answer

Authentication establishes a caller’s identity. Authorization checks a specific action on a specific resource. Tenant isolation requires those checks and data boundaries to prevent cross-tenant exposure through every path. For example, authenticated principal U10 belongs to Birch and must not read Acme invoice I17 through the API, cache, search, export, or file endpoint.

What the answer must demonstrate: Use an actual permitted and forbidden resource path.

Can the API trust an X-Tenant-ID header?

Reveal a model answer

“It can treat it as a requested tenant, then verify the authenticated principal’s membership and permission. I never let a caller-selected tenant ID bypass that decision, and I propagate the verified scope into data access.”

What the answer must demonstrate: Trace how the scope becomes trusted.

Why can invoice:I17 be an unsafe cache key?

Reveal a model answer

Two tenants may share invoice I17, and users may have different field permissions. I key cached bodies by tenant and immutable representation version, authorize the exact version and current policy scope, then return only that authorized representation. If the fetched body or required policy revision differs from the decision, I reauthorize or withhold it.

What the answer must demonstrate: A fast cache can consistently leak data.

Is row-level security sufficient on its own?

Reveal a model answer

“It is useful defense in depth when policies, roles, and connection context are correct. I still enforce object/action permission in the application and verify privileged-role bypass behavior. A database policy cannot secure an unscoped object-storage or cache path.”

What the answer must demonstrate: Know the enforcement boundary.

Does encryption at rest prevent one tenant seeing another’s records?

Reveal a model answer

“No. If the application can decrypt both tenants’ records, it can still send the wrong one. Check tenant permissions, route to the correct data and return only allowed fields. Encryption protects stored bytes; it does not make those application decisions.”

What the answer must demonstrate: Name the threat each mechanism addresses.

How do you stop one tenant’s exports slowing every customer?

Reveal a model answer

“I bound per-tenant concurrency and total queues, schedule fairly, and separate heavy export workers from interactive reads. Quotas describe an enforceable budget; admission control prevents accepting more work than we can serve.”

What the answer must demonstrate: Security includes resource isolation.

What would you test beyond successful login?

Reveal a model answer

“Use two tenants and attempt cross-tenant reads, writes, search, exports, attachment downloads, and cache hits. Also test revoked membership and expired capabilities. Each denied operation must leave data and side effects unchanged under its contract.”

What the answer must demonstrate: Exercise alternate access paths.

What are RPO and RTO in disaster recovery?

Reveal a model answer

RPO, recovery point objective, is the targeted maximum data-loss window. RTO, recovery time objective, is the targeted time to restore the agreed service. If the last recoverable state is 5 seconds before a disruption and service returns 7 minutes later, those are separate data-loss and restoration measurements to compare with the objectives.

What the answer must demonstrate: Keep the two objectives separate and distinguish targets from measured guarantees.

Can asynchronous regional replication promise zero loss of acknowledged writes?

Reveal a model answer

“Not by itself. East can acknowledge a write and fail before it reaches West. To survive that regional loss without losing acknowledged writes, acknowledgement must require a durable copy or quorum outside East, within the stated failure model.”

What the answer must demonstrate: Place the acknowledgement boundary.

The East primary stops responding to West. Why is that alone insufficient to promote West safely?

Reveal a model answer

A failed health check cannot prove that East stopped writing. In the asynchronous two-region design, I require a verified stop or removal of its write capability before promotion; if that is impossible, writes remain paused. Alternatively, a proven quorum protocol prevents the isolated minority from committing. A new epoch stored only in West is not sufficient fencing.

What the answer must demonstrate: Separate routing from write authority.

Would active-active remove all regional outages?

Reveal a model answer

“It can improve continuity for some operations, but shared dependencies and write conflicts remain. I would state whether each key has one home writer, uses global coordination, or permits a defined merge. Inventory cannot simply merge arbitrary decrements without a rule.”

What the answer must demonstrate: Describe per-record semantics.

Can a nightly backup meet a 30-second RPO?

Reveal a model answer

“A snapshot alone cannot: it can leave almost a day of changes absent. Continuous recoverable logs or another replication mechanism may narrow that gap. I also need to test restore and replay time against the RTO.”

What the answer must demonstrate: Check both freshness and duration.

The standby has one fifth of peak capacity. Is failover ready?

Reveal a model answer

“Only if the recovery contract allows bounded degradation and the remaining capacity or scaling is verified. I would test databases and dependencies too, prioritize essential operations, and limit admission rather than overload the new primary.”

What the answer must demonstrate: Capacity is part of recovery.

Why keep backups when there are three replicas?

Reveal a model answer

“Replicas can copy an accidental deletion or corruption. Backups and point-in-time recovery preserve earlier states under a separate protection policy. I would regularly restore and validate business records, not only check the backup job status.”

What the answer must demonstrate: Replication is not historical recovery.

An old primary region recovers after failover. Why should writes not immediately be routed back?

Reveal a model answer

“West has accepted new writes, so keep it in charge. Bring East up to date or rebuild it from West, verify the data, then switch writers through a controlled handover. Prevent the former writer from continuing. If the regions have conflicting histories, resolve them before switching.”

What the answer must demonstrate: Failback is a controlled state transition.

Define SLI, SLO, SLA, and error budget with one user-visible example.

Reveal a model answer

An SLI (service-level indicator) is a measured user-visible behavior. An SLO is its target over a window. An SLA is an agreement about service commitments and consequences, often contractual. An error budget is the amount of failure permitted by the SLO.

For a photo service, measure the fraction of eligible accepted photos that become ready within 60 seconds, after each photo's evaluation period has elapsed. Set an illustrative SLO of at least 99.9% over 30 days. Of one million evaluated photos, at most 1,000 can miss the target. A separate SLA could specify a contractual commitment and credits; it need not use the same threshold as the internal SLO. Also measure acceptance so the service cannot make its completion ratio look good by rejecting every upload.

What the answer must demonstrate: A percentage without a denominator and window is incomplete.

Could accepting no uploads make your completion SLO look perfect?

Reveal a model answer

“Yes, if it treats no data as success, or shows only accepted uploads and hides rejected attempts. Zero evaluated photos proves nothing about completion. Also measure how many valid attempts are accepted, signal missing data and use a test upload when useful. Then we can distinguish failure to accept uploads from failure to process them.”

What the answer must demonstrate: Beware metrics that improve by refusing useful work.

For 1,000,000 evaluated operations and a 99.9% success SLO, what is the error allowance? What burn rate does a 2% bad-event rate represent?

Reveal a model answer

“At 99.9%, one million evaluated operations allow one thousand bad events. If a recent window has two percent bad events against a 0.1 percent allowance, its burn rate is twenty. I would interpret that with traffic and window size before deciding how urgently to page.”

What the answer must demonstrate: Keep percentage points and ratios distinct.

A photo takes 100 seconds: 95 queued, 4 rendering, 1 publishing. What does this trace reveal that low CPU usage does not?

Reveal a model answer

“The trace assigns ninety-five seconds to waiting, four to rendering, and one to publication. Low CPU cannot tell me whether concurrency is accidentally restricted or demand is absent. The trace locates the delay; release/configuration evidence then helps identify the cause.”

What the answer must demonstrate: Separate symptom, location, and causal evidence.

When do you use metrics, logs, and traces?

Reveal a model answer

“Metrics show aggregate trends and support alerts. Structured logs record individual events. Traces or correlated job events connect a specific journey across stages. For P501 I use metrics to detect late completion and the trace plus targeted logs to explain where it waited and which release handled it.”

What the answer must demonstrate: Choose the evidence type according to the question.

What should the canary compare before full deployment?

Reveal a model answer

“Compare similar workloads, user-visible completion, throughput and stage delay, with enough observations to distinguish a signal from noise. For worker changes, separate canary and control pools can make queue-wait attribution meaningful. If both versions share a queue, rising age affects the cohort comparison and cannot by itself blame one version. I would inspect per-version throughput/configuration and stop expansion or revert when the canary violates the agreed guardrails.”

What the answer must demonstrate: Deployment safety includes data compatibility.

The old worker version is back. Can you close the incident?

Reveal a model answer

“Only after verifying the backlog drains and timely completion recovers. I also check that retrying a job did not publish duplicate results or repeat other side effects, and that photo permissions remain correct. Restoring the old version is a recovery step; I still need to verify that users can upload and view photos successfully.”

What the answer must demonstrate: Verify recovery under continuing load.

How do RPO and RTO change your recovery exercise?

Reveal a model answer

“RPO tells me how much recent accepted work may be lost; RTO tells me how soon useful service should return. I would measure both during a drill and verify original files, records of pending jobs, publication status, and permissions, rather than timing only a database restore command.”

What the answer must demonstrate: Recovery objectives apply to the service outcome.

Foundation · Question 1Design a URL shortener

Does the shortener download the destination page?

Reveal a model answer

No. It returns a redirect response containing the destination URL. The browser then contacts that website separately.

What the answer must demonstrate: Separates the redirect response from the browser’s destination request.

Applied · Question 2Design a URL shortener

Why is a large random code space not enough?

Reveal a model answer

Two requests can still generate the same candidate. The database’s unique constraint accepts only one competing insert; the loser chooses another code.

What the answer must demonstrate: Names an atomic uniqueness constraint and the check-then-write race.

Applied · Question 3Design a URL shortener

How does a retry avoid creating a second link?

Reveal a model answer

Scope a request key to the authenticated owner and store its payload identity and result atomically with the mapping. Replay a matching completed request.

What the answer must demonstrate: Keeps request identity, payload validation and mapping commit connected.

Applied · Question 4Design a URL shortener

Why can caching change deletion behavior?

Reveal a model answer

The database can contain a deletion while a cache still holds the prior mapping. Serving the copy therefore needs an agreed freshness policy.

What the answer must demonstrate: States a deletion visibility contract instead of treating invalidation as guaranteed.

Applied · Question 5Design a URL shortener

Why does hashing codes not fix one viral link?

Reveal a model answer

Hashing distributes different codes. Every lookup for the same code still targets its owner.

What the answer must demonstrate: Distinguishes distribution of different keys from repeated reads of one key.

Foundation · Question 6Design a URL shortener

Can a delayed cleanup job extend a link lifetime?

Reveal a model answer

No. Every serving path checks expiry. Cleanup removes obsolete storage later.

What the answer must demonstrate: Separates expiry enforcement from physical deletion and permanent claims.

Follow-up · Question 7Design a URL shortener

What must be reconsidered after splitting storage?

Reveal a model answer

The mapping and request result may no longer share one local transaction. Retain transactional support or specify a recoverable creation workflow.

What the answer must demonstrate: Recognizes which transaction boundary the new storage layout removes.

Follow-up · Question 8Design a URL shortener

What changes for private links?

Reveal a model answer

Authenticate readers and evaluate current access before returning a destination; code secrecy does not establish permission.

What the answer must demonstrate: Separates possession of a URL from current reader authorization.

Foundation · Question 1Design Pastebin

What changes compared with a URL shortener?

Reveal a model answer

The service stores and serves the actual text, so completeness, byte transfer, safe rendering and content permissions become central.

What the answer must demonstrate: Names actual content ownership rather than treating the product as another redirect.

Foundation · Question 2Design Pastebin

Is an unlisted paste private?

Reveal a model answer

No. Anyone with its address can retrieve it. Private content requires an authenticated grant.

What the answer must demonstrate: Distinguishes access by URL possession from authenticated access.

Applied · Question 3Design Pastebin

Why begin with text in SQL?

Reveal a model answer

At the assumed request rate it provides one simple transaction for text, metadata and retry state.

What the answer must demonstrate: Uses operational pressure to justify losing a one-store transaction.

Applied · Question 4Design Pastebin

Why upload the object before marking READY?

Reveal a model answer

A failed upload must not leave a visible paste pointing to missing bytes. Hidden orphan bytes are safer and recoverable.

What the answer must demonstrate: Explains the safe ordering and the orphan-versus-broken-pointer tradeoff.

Applied · Question 5Design Pastebin

Why is an old upload not automatically safe to delete?

Reveal a model answer

Its uploader may still be completing. Cleanup must first mark the attempt canceled in the same database state that publication checks, then delete its bytes.

What the answer must demonstrate: Identifies a state guard shared by publication and deletion.

Applied · Question 6Design Pastebin

Can we cache private paste bodies?

Reveal a model answer

Internally, yes, provided every new delivery passes the required current permission check.

What the answer must demonstrate: Keeps permission enforcement in front of every byte-delivery path.

Applied · Question 7Design Pastebin

Why limit bytes as well as requests?

Reveal a model answer

A maximum-size paste consumes far more bandwidth and buffering than an average one. Request counts alone hide that imbalance.

What the answer must demonstrate: Connects object size, transfer time and bounded resources.

Follow-up · Question 8Design Pastebin

How would editing change the design?

Reveal a model answer

Store immutable body versions and atomically change the selected version using an expected-version check.

What the answer must demonstrate: Separates immutable content versions from mutable selection metadata.

Foundation · Question 1Design a photo-sharing service

When is an uploaded photo ready for a feed?

Reveal a model answer

After every required display variant exists and its accepted manifest is committed.

What the answer must demonstrate: Separates original acceptance from published variants.

Applied · Question 2Design a photo-sharing service

Why store processing work beside photo state?

Reveal a model answer

A durable outbox committed with PROCESSING avoids losing the job if the application crashes before a separate queue send.

What the answer must demonstrate: Explains the database-to-queue crash gap and repeated relay.

Applied · Question 3Design a photo-sharing service

Why not push every photo to every follower?

Reveal a model answer

Popular authors can create millions of writes for inactive readers. Pulling their recent timelines on demand can be cheaper.

What the answer must demonstrate: Justifies the threshold with work, rather than a celebrity label alone.

Applied · Question 4Design a photo-sharing service

Why check access after reading a feed cache?

Reveal a model answer

Candidate lists can outlive an unfollow, deletion or privacy change. Membership in that list is not permission.

What the answer must demonstrate: Keeps authorization independent of derived-list cleanup.

Applied · Question 5Design a photo-sharing service

How does a stale image worker fail safely?

Reveal a model answer

It writes attempt-specific outputs and must pass a current-attempt check before publishing the manifest.

What the answer must demonstrate: Protects both metadata selection and immutable bytes.

Applied · Question 6Design a photo-sharing service

Does a CDN eliminate image delivery cost?

Reveal a model answer

It reduces origin reads and can improve latency, but bytes delivered to viewers still consume bandwidth and cost money.

What the answer must demonstrate: Separates origin savings from total viewer egress.

Follow-up · Question 8Design a photo-sharing service

What changes for a personalized feed?

Reveal a model answer

Add bounded candidate scoring and evaluate quality, while preserving publication and permission checks.

What the answer must demonstrate: Treats relevance and eligibility as separate decisions.

Why separate file ID from path?

Reveal a model answer

A rename changes the directory entry, not which file and history the devices are tracking.

What the answer must demonstrate: Distinguishes stable identity from mutable directory placement.

What happens when two devices edit revision 12?

Reveal a model answer

The first valid commit changes the head. The second expected-revision check fails and preserves its work as a conflict.

What the answer must demonstrate: Uses an atomic expected-version check and preserves the losing edit.

What does chunking save in the 9 MiB example?

Reveal a model answer

Changing only the middle 4 MiB allows the other 5 MiB to be reused.

What the answer must demonstrate: Calculates changed bytes without promising universal deduplication.

Why is a notification not enough for synchronization?

Reveal a model answer

Notifications can disappear while a device is offline. A durable change log and cursor recover what was missed.

What the answer must demonstrate: Treats cursor recovery as durable state rather than transport behavior.

When does an upload become the current file?

Reveal a model answer

The server verifies that chunks exist and cannot be deleted, checks current permission and the expected base revision, then commits the new revision in metadata.

What the answer must demonstrate: Separates byte upload from publication and request replay.

Why keep old-revision references during cleanup?

Reveal a model answer

Historical versions still need their chunks for restoration. Checking only which chunks the current revision uses would lose older versions.

What the answer must demonstrate: Counts retained history and active-upload protection.

What can fail on the receiving device?

Reveal a model answer

A crash can occur between file replacement and local cursor persistence, so a recovery journal must connect them.

What the answer must demonstrate: Recognizes the local filesystem/metadata crash boundary.

Can the service automatically merge spreadsheets?

Reveal a model answer

Only with format-specific semantics and a defined conflict policy; a generic byte synchronizer cannot infer user intent.

What the answer must demonstrate: Avoids inventing merge semantics for arbitrary binary formats.

Foundation · Question 1Design a chat messaging service

What does accepted mean?

Reveal a model answer

The message has been committed under the storage durability policy, independently of recipient connectivity.

What the answer must demonstrate: Keeps durability, device receipt and reading as separate outcomes.

Foundation · Question 2Design a chat messaging service

Does WebSocket make chat reliable?

Reveal a model answer

It provides a bidirectional transport, not durable message storage or retry semantics.

What the answer must demonstrate: Attributes reliability to storage and recovery rather than a transport name.

Applied · Question 3Design a chat messaging service

Why deduplicate by identity rather than text?

Reveal a model answer

Identical text can be sent intentionally twice. The stable send identity distinguishes a retry from a new message.

What the answer must demonstrate: Separates intentional repeated content from repeated transmission.

Applied · Question 4Design a chat messaging service

Why use a conversation sequence?

Reveal a model answer

Devices can disagree about network arrival and client clocks. One committed sequence provides their shared history order.

What the answer must demonstrate: Uses local ordering without unnecessary global coordination.

Applied · Question 5Design a chat messaging service

What if sequence 1042 arrives after cursor 1040?

Reveal a model answer

Fetch the missing authorized history before advancing a contiguous delivered cursor.

What the answer must demonstrate: Detects gaps and persists device progress safely.

Applied · Question 6Design a chat messaging service

What survives a gateway crash?

Reveal a model answer

Committed messages and device recovery identities survive in storage; clients reconnect and catch up.

What the answer must demonstrate: Distinguishes replaceable connections from authoritative message state.

Applied · Question 7Design a chat messaging service

Why can groups cost more than their stored messages?

Reveal a model answer

One message may be delivered to many members and devices, creating fanout work beyond one database insert.

What the answer must demonstrate: Counts recipient amplification and bounds work.

Follow-up · Question 8Design a chat messaging service

What changes with end-to-end encryption?

Reveal a model answer

Clients manage message-encryption keys, and servers store ciphertext rather than freely processing content.

What the answer must demonstrate: Recognizes that encryption changes product and recovery semantics.

Foundation · Question 1Design a microblogging service

How can the first version build a home feed?

Reveal a model answer

Read followed authors, query recent author-indexed posts, merge by time and ID, then filter visibility. This is complete without an inbox service.

What the answer must demonstrate: Explains the complete pull path and its repeated work.

Applied · Question 2Design a microblogging service

Why not push every post to every follower?

Reveal a model answer

A celebrity can require millions of recipient writes for one post, including inactive viewers. Keep those author histories on a read-time merge path.

What the answer must demonstrate: Connects follower skew to hybrid fanout, not a universal threshold.

Applied · Question 3Design a microblogging service

Why save an outbox event with a post?

Reveal a model answer

A crash between database commit and queue publication would otherwise lose the fanout trigger. The saved event can be relayed again.

What the answer must demonstrate: Identifies the commit-to-queue gap and repeatable consumers.

Applied · Question 4Design a microblogging service

A worker crashes halfway through a follower page. What happens?

Reveal a model answer

Repeat the page; unique viewer/post keys suppress duplicate effects. Advance progress only after all writes complete.

What the answer must demonstrate: Places checkpoint after effects and uses unique inbox keys.

Applied · Question 5Design a microblogging service

Why does deleting every inbox entry not suffice for privacy?

Reveal a model answer

Cleanup may be delayed or incomplete. The read service must consult current post visibility and private membership before disclosing content.

What the answer must demonstrate: Separates candidate freshness from authorization.

Follow-up · Question 6Design a microblogging service

Does a keyset cursor create a frozen feed?

Reveal a model answer

No. It stabilizes the continuation position, but late fanout can add older candidates. The simple contract allows those items to appear on refresh.

What the answer must demonstrate: States the keyset limitation before adding a session snapshot.

Applied · Question 7Design a microblogging service

Why store likes as relationships?

Reveal a model answer

A unique user/post pair makes repeated like and unlike requests well-defined. Blind counter increments duplicate actions after retries.

What the answer must demonstrate: Distinguishes authoritative relationship from derived count.

Applied · Question 8Design a microblogging service

Why distinguish page requests from item impressions?

Reveal a model answer

A page contains many records, media references and visibility checks. Multiplying by items per page exposes the actual serving work.

What the answer must demonstrate: Keeps page, item and fanout units separate.

Foundation · Question 1Design a video streaming service

Why is upload completion different from READY?

Reveal a model answer

The original can be durable while no compatible playback output exists. READY requires a verified selected manifest and required files.

What the answer must demonstrate: Distinguishes durable bytes from verified playable outputs.

How do you estimate concurrent viewers?

Reveal a model answer

Multiply playback starts per second by average watched seconds, then multiply concurrency by average selected bitrate for delivery throughput.

What the answer must demonstrate: Keeps starts, concurrency, bits and bytes in correct units.

Why does a lower rendition help on a slow connection?

Reveal a model answer

It needs fewer bytes for the same media duration, allowing downloads to replenish the buffer faster than playback consumes it.

What the answer must demonstrate: Explains buffer supply versus consumption causally.

How can a successfully loaded manifest still produce broken playback?

Reveal a model answer

It may name segments that were never written or were overwritten. Verify all required immutable outputs before selecting the manifest.

What the answer must demonstrate: Protects both manifest selection and the referenced bytes.

What does a CDN solve and not solve?

Reveal a model answer

It reuses bytes near viewers and reduces origin load. It does not create source durability or authorize private content by itself.

What the answer must demonstrate: Separates cache reuse from durability and authorization.

Follow-up · Question 6Design a video streaming service

Can a five-minute playback grant support immediate revocation?

Reveal a model answer

No. An already issued grant can remain usable until expiry; current authorization is required for a new one.

What the answer must demonstrate: States the actual grant-expiry limitation.

Why not save the largest playback offset ever received?

Reveal a model answer

A person may intentionally seek backward, and old devices may send late updates. Larger time is not necessarily newer intent.

What the answer must demonstrate: Orders user intent by session/events rather than maximum offset.

Which metrics reveal whether playback works?

Reveal a model answer

First-frame delay, rebuffered fraction, device/codec errors and actual decoded segment checks reveal user experience better than metadata response codes.

What the answer must demonstrate: Measures decoded user experience and explicit cost tradeoffs.

Foundation · Question 1Design a typeahead service

How does a trie answer a prefix query?

Reveal a model answer

Follow character edges to the prefix node, then use descendant terminal terms. A terminal marker distinguishes a complete term from an intermediate path.

What the answer must demonstrate: Explains prefix traversal and terminal semantics.

Applied · Question 2Design a typeahead service

Why is prefix-length traversal not the whole query cost?

Reveal a model answer

A broad prefix can have millions of matching descendants. Enumerating and sorting them dominates the few edge traversals.

What the answer must demonstrate: Counts descendant enumeration, not just edge traversal.

Applied · Question 3Design a typeahead service

What goes wrong when a top-ranked term loses popularity?

Reveal a model answer

The old shortlist does not contain the best excluded replacement. Retained vocabulary and counts are needed to recompute winners.

What the answer must demonstrate: Finds the missing replacement candidate after a decrease.

Applied · Question 4Design a typeahead service

Why build immutable snapshots?

Reveal a model answer

Build the replacement separately so a query never combines a changed term score with an old prefix shortlist. Each request uses one complete version of the terms, normalization rules and scores.

What the answer must demonstrate: States coherent data and simultaneous-version memory costs.

Applied · Question 5Design a typeahead service

Why does request cancellation not prevent stale display?

Reveal a model answer

The old response may already be completed or in flight. The browser must compare its input sequence before rendering.

What the answer must demonstrate: Uses an explicit latest-input comparison.

Applied · Question 6Design a typeahead service

Why prefer lexical ranges to hashing complete terms here?

Reveal a model answer

Prefix matches occupy related lexical ranges, allowing targeted routing. Hashing terms scatters matches and usually requires broad fanout.

What the answer must demonstrate: Connects partition layout to query locality and hot-key replication.

Applied · Question 7Design a typeahead service

Where should blocked-term filtering run?

Reveal a model answer

Combine public index/cache candidates with the authenticated user’s private-history candidates, then check the whole list before returning it.

What the answer must demonstrate: Checks the combined candidate list and does not refill it from an unchecked source.

Follow-up · Question 8Design a typeahead service

Why is a personalized result unsafe in a public cache?

Reveal a model answer

It may reveal one user’s history to another. Share only public candidates and perform authenticated personal merging separately.

What the answer must demonstrate: Keeps personal data out of shared results and recognizes candidate recall limits.

Foundation · Question 1Design an API rate limiter

What must be clarified before choosing a limiter algorithm?

Reveal a model answer

Ask whose allowance is shared, which action consumes it, over what interval, whether bursts are allowed and whether an outage should admit or deny. Those answers determine the algorithm and stored state.

What the answer must demonstrate: States all dimensions of the product contract before storage.

Applied · Question 2Design an API rate limiter

Why does a fixed-minute counter violate this rolling policy?

Reveal a model answer

It can admit three just before a clock boundary and three immediately after, all inside one rolling minute.

What the answer must demonstrate: Demonstrates the boundary counterexample and distinguishes burst semantics.

Applied · Question 3Design an API rate limiter

Two gateways see two used slots. How do you avoid admitting both?

Reveal a model answer

Serialize prune, count, compare and insertion inside one quota-key transaction. The second transaction sees the first accepted event.

What the answer must demonstrate: Makes the eligibility check and effect one atomic operation.

Applied · Question 4Design an API rate limiter

What happens when an allow reply disappears?

Reveal a model answer

Retry the trusted internal decision ID and recover its saved answer without another usage event.

What the answer must demonstrate: Separates internal decision recovery from business idempotency.

Follow-up · Question 5Design an API rate limiter

Why is asynchronous counter replication insufficient for strict failover?

Reveal a model answer

A promoted replica may lack an acknowledged admission and grant extra allowance. The selected datastore must preserve committed usage or stop admission.

What the answer must demonstrate: Connects acknowledged-state survival to the chosen availability tradeoff.

Applied · Question 6Design an API rate limiter

Why can you cache a deny hint but not an allow?

Reveal a model answer

A conservative denial cannot add extra admissions; reusing an allow would skip the state-changing consume operation.

What the answer must demonstrate: Uses denial conservatively and never turns cached permission into free quota.

Applied · Question 7Design an API rate limiter

Can a rate limit alone protect expensive long-running work?

Reveal a model answer

No. Arrival rate multiplied by execution duration determines in-flight work, so a separate concurrency limit or bounded queue may be needed.

What the answer must demonstrate: Separates arrival rate from in-flight resource use.

Applied · Question 8Design an API rate limiter

What failure can a forward clock jump cause?

Reveal a model answer

It can make recent admissions appear older than the window and release allowance too early.

What the answer must demonstrate: Explains premature expiry and the explicit trusted-time limit.

Foundation · Question 1Design public post search

How does solar AND battery find its matches?

Reveal a model answer

Intersect the solar and battery posting lists, yielding T101 and T103 in the example. OR would union IDs and deduplicate.

What the answer must demonstrate: Demonstrates intersection/union before ordering.

Applied · Question 2Design public post search

Why can a saved post be missing from search?

Reveal a model answer

Source commit and index refresh are separate stages. A durable change event allows recovery while search visibility is temporarily delayed.

What the answer must demonstrate: Separates committed source, recoverable event and searchable refresh.

Applied · Question 3Design public post search

Why choose time/document shards?

Reveal a model answer

Each post’s terms update together, each shard evaluates the full Boolean query, and time filters skip old buckets.

What the answer must demonstrate: Connects document locality and time pruning to fanout cost.

Applied · Question 4Design public post search

How does a delete defeat a late edit?

Reveal a model answer

Store the source-assigned latest version and atomically reject older updates. Keep the delete version through the replay horizon.

What the answer must demonstrate: Uses an atomic version comparison and retains deletion evidence.

Applied · Question 5Design public post search

When may an indexer advance its source position?

Reveal a model answer

After the indexed changes are recoverable. Replaying already applied versions is safe; skipping undurable changes is not.

What the answer must demonstrate: Never lets progress outrun recoverable index state.

Applied · Question 6Design public post search

Why combine a snapshot and search-after?

Reveal a model answer

The snapshot fixes the visible index state, and the last sort tuple gives a stable continuation position without shifting offsets.

What the answer must demonstrate: Combines stable ordering with current deletion checks.

Applied · Question 7Design public post search

An indexed match now has different text. What should be returned?

Reveal a model answer

Authorize and load the same current version, then verify it still satisfies the query. Omit a nonmatch; the index may temporarily miss new matches too.

What the answer must demonstrate: Binds content and permission versions and acknowledges recall lag.

Follow-up · Question 8Design public post search

Why is reranking twenty recent matches not globally most-liked search?

Reveal a model answer

A highly liked older match may be absent from the retrieved set. Popularity must affect candidate selection or the result must be labeled approximate.

What the answer must demonstrate: Recognizes missing candidates and score comparability requirements.

Foundation · Question 1Design a web crawler

What is the frontier and why must it be durable?

Reveal a model answer

It records URLs waiting for initial fetch or revisit. Losing it loses accepted future work, even if previously fetched bodies survive.

What the answer must demonstrate: Distinguishes future work from completed bodies and discovery.

Applied · Question 2Design a web crawler

Why do per-worker delays fail to enforce politeness?

Reveal a model answer

Multiple workers can each obey their own delay while sending concurrent requests to the same origin. Origin-wide state must coordinate starts and active requests.

What the answer must demonstrate: Coordinates all workers at the origin and respects site-bound capacity.

Applied · Question 3Design a web crawler

Can a crawler promise exactly-once HTTP fetches?

Reveal a model answer

No. A timeout can leave the request outcome unknown, and recovery may repeat the network operation. Durable idempotent processing is the defensible promise.

What the answer must demonstrate: Allows duplicate network attempts while guarding state updates.

Applied · Question 4Design a web crawler

How do you recover after inserting U18 but before U19?

Reveal a model answer

Replay the durable manifest. Exact unique URL keys make U18 harmless to repeat, then U19 is inserted before progress advances.

What the answer must demonstrate: Keeps manifest progress behind durable exact URL insertion.

Applied · Question 5Design a web crawler

Why must a redirect be checked again?

Reveal a model answer

Its target may be outside scope or resolve to a private network even when the original address was public.

What the answer must demonstrate: Validates actual connection destinations and every redirect.

Applied · Question 6Design a web crawler

Why cannot a Bloom positive prove a URL was seen?

Reveal a model answer

False positives can mark an unseen URL as possibly present. Discarding it without exact lookup silently loses coverage.

What the answer must demonstrate: Identifies false-positive recall loss and exact-store authority.

Applied · Question 7Design a web crawler

Can identical HTML skip all link extraction work?

Reveal a model answer

No. Relative links resolve using each page’s effective URL, so the same bytes on two origins can discover different addresses.

What the answer must demonstrate: Preserves effective-URL context when sharing bytes.

Applied · Question 8Design a web crawler

What should happen when parsing falls behind?

Reveal a model answer

Keep fetched bodies durable and reduce new intake until the downstream stage catches up. Unbounded downloading only moves the failure to storage.

What the answer must demonstrate: Uses backpressure and useful-content outcomes rather than raw throughput.

Foundation · Question 1Design a personalized news feed

What are the four responsibilities in a feed?

Reveal a model answer

Publication stores the source, candidate generation finds possible stories, ranking orders them and delivery returns selected authorized content.

What the answer must demonstrate: Keeps durability and notifications separate.

Applied · Question 2Design a personalized news feed

Why combine write-time and read-time generation?

Reveal a model answer

Ordinary active audiences benefit from prepared references, while celebrity fanout may create millions of unused writes. Pull those histories during actual reads.

What the answer must demonstrate: Uses audience economics to justify both paths.

Applied · Question 3Design a personalized news feed

Why not represent every relationship as a follow?

Reveal a model answer

Distribution interest, approved friendship and private-group membership grant different behavior and access. A one-way follow cannot silently authorize friends-only content.

What the answer must demonstrate: Distinguishes distribution from actual access grants.

Applied · Question 4Design a personalized news feed

What if fanout inserts a story after the viewer unfollows?

Reveal a model answer

The late candidate can remain temporarily, but current serving-time relationship checks exclude it. Async cleanup is an optimization.

What the answer must demonstrate: Checks current eligibility and binds the returned body version.

Applied · Question 5Design a personalized news feed

How can a ranked feed keep page two stable?

Reveal a model answer

Save a bounded ordered candidate session and resume by its cursor, while rechecking current eligibility on each page.

What the answer must demonstrate: Freezes bounded order without freezing permissions.

Applied · Question 6Design a personalized news feed

What is the difference between ranking quality and candidate recall?

Reveal a model answer

A ranker can order only retrieved candidates. A useful story excluded by an overly small pool cannot be recovered by a perfect score.

What the answer must demonstrate: Separates missing candidates from scoring errors.

Applied · Question 7Design a personalized news feed

How do you avoid a gap while rebuilding a lost candidate list?

Reveal a model answer

Record an event boundary before scanning source timelines, then replay overlapping changes and deduplicate before cutover.

What the answer must demonstrate: Includes overlap/replay rather than scan-only reconstruction.

Applied · Question 8Design a personalized news feed

What can safely degrade during a ranker outage?

Reveal a model answer

Use chronological ordering over bounded eligible candidates. Current permissions and content-version checks remain mandatory.

What the answer must demonstrate: Degrades ordering without weakening authorization.

Why can a café across a cell boundary be the closest result?

Reveal a model answer

Cells organize storage; they do not constrain physical distance. Maya at x=990 and P12 at x=1010 are twenty meters apart. Search must cover intersecting cells and apply exact distance.

What the answer must demonstrate: Explain cross-cell coverage and the exact circular distance predicate.

When can a nearest-twenty query stop?

Reveal a model answer

After examining all eligible bounded candidates, or after every remaining region has a valid minimum distance worse than the current twentieth result. Equal distances need the stated tie-breaker.

What the answer must demonstrate: State the stopping bound and apply eligibility before counting best results.

Why name the field radiusMeters?

Reveal a model answer

It makes the distance unit explicit. Latitude and longitude are angles, and longitude distance varies with latitude, so raw degree subtraction is unsuitable.

What the answer must demonstrate: Distinguish angular coordinates from distance units and use appropriate geography.

Why partition spatially instead of hashing place IDs?

Reveal a model answer

Geographic ownership keeps most local queries on a few owners. Hashing IDs balances records but requires querying every index partition.

What the answer must demonstrate: Compare geographic locality with ID-hash fanout and hot-city skew.

Can current detail reads repair every stale-index error?

Reveal a model answer

They can remove incorrect returned candidates but cannot discover a moved place absent from the new region. Preserve the indexing-delay contract or use a caught-up index.

What the answer must demonstrate: Distinguish rejecting stale candidates from discovering missing moved records.

What is the simplest nearby-friends design?

Reveal a model answer

Fetch the viewer’s authorized contacts, read their latest positions in a batch, check expiry and compute distances. A bounded contact set may not need a global spatial index.

What the answer must demonstrate: Use current viewer consent and presence expiry, including the small-contact-set alternative.

A neighboring region is unavailable; can twenty other results count as success?

Reveal a model answer

Not for an exact nearest-result promise. The missing region might contain closer places. Return an explicitly incomplete result or fail that query.

What the answer must demonstrate: Refuse to claim exact completeness when a required region is missing.

Is best-rated within a radius the same as nearest twenty?

Reveal a model answer

No. A farther café inside the radius may outrank every nearby café on rating. Candidate selection must cover the ranking contract.

What the answer must demonstrate: Choose candidate coverage from the ranking objective, not only nearest distance.

Foundation · Question 1Design a ride-hailing backend

Why separate driver position from availability?

Reveal a model answer

Position describes an observation; availability is a durable promise about whether the driver can accept work. A fresh heartbeat must not release an active assignment.

What the answer must demonstrate: Separate observed location from authoritative assignment availability.

Applied · Question 2Design a ride-hailing backend

D17 accepts R501 and R502 concurrently. What prevents two riders winning?

Reveal a model answer

Both transactions lock or conditionally guard the same driver record and atomically update the corresponding ride. The second sees D17 assigned and changes neither side.

What the answer must demonstrate: Protect both driver exclusivity and ride exclusivity in the same transaction.

Applied · Question 3Design a ride-hailing backend

The acceptance reply is lost. What should the retry return?

Reveal a model answer

The saved assignment for the same authenticated operation or accepted offer. Check that result after lock acquisition so a concurrent duplicate observes the committed winner.

What the answer must demonstrate: Recover the saved assignment after lock acquisition instead of rejecting a successful retry.

Foundation · Question 4Design a ride-hailing backend

Why include a session generation as well as a sequence?

Reveal a model answer

Sequences order updates within one publishing session. A generation distinguishes a restarted or replacement session and fences delayed packets from the old one.

What the answer must demonstrate: Explain session replacement, sequence ordering and observation-time limits.

Applied · Question 5Design a ride-hailing backend

Why send only a few offers at a time?

Reveal a model answer

Bounded batches limit driver interruption and expensive routing work. Each batch rechecks that the ride is still eligible.

What the answer must demonstrate: Connect bounded candidate work to driver interruption and matching latency.

Foundation · Question 6Design a ride-hailing backend

Should a lost socket make a driver available?

Reveal a model answer

No. The current trip is durable. The phone reconnects, authenticates and recovers its assignment; stale location only affects tracking and discovery.

What the answer must demonstrate: Preserve durable trip ownership across socket loss and distinguish coordinate buffering.

Follow-up · Question 7Design a ride-hailing backend

How does cancellation race with acceptance?

Reveal a model answer

Both use the ride authority. Cancellation first blocks acceptance; assignment first means cancellation follows the assigned-trip policy and releases both references atomically if allowed.

What the answer must demonstrate: Serialize cancellation with acceptance and use committed versions for notifications.

Follow-up · Question 8Design a ride-hailing backend

Why not change assignment owner on every spatial-cell crossing?

Reveal a model answer

Cells help find nearby drivers. Assignment ownership keeps ride and driver records consistent. Moving those records on every cell crossing would add unnecessary coordination.

What the answer must demonstrate: Distinguish spatial movement from regional assignment-authority transfer.

Foundation · Question 1Design a ticket-booking service

Why distinguish a hold from a row lock?

Reveal a model answer

A hold is a durable reservation with a business deadline. A row lock protects a short state transition; holding it while a person pays wastes connections and does not define recovery.

What the answer must demonstrate: Distinguish minutes-long business reservations from short database locks.

Applied · Question 2Design a ticket-booking service

How do seats 54–56 and 56–57 avoid a partial allocation?

Reveal a model answer

Each transaction locks its selected seats in order, validates the full set and commits all changes together. The loser on seat 56 rolls back its entire set.

What the answer must demonstrate: Validate and commit the entire requested seat set or roll it back.

Applied · Question 3Design a ticket-booking service

The hold response disappears after commit. What next?

Reveal a model answer

Retry the same scoped key and payload to retrieve the saved hold. Claim that key before allocation so concurrent duplicates serialize correctly.

What the answer must demonstrate: Claim a scoped request identity and compare payloads before allocating seats.

Applied · Question 4Design a ticket-booking service

A payment times out. Can the customer be charged again?

Reveal a model answer

Not merely because of the timeout. Reuse the stored attempt identity and reconcile its outcome under the provider contract.

What the answer must demonstrate: Persist a stable attempt before the provider call and retain unknown outcomes.

Applied · Question 5Design a ticket-booking service

Payment success arrives after another customer holds the seats. What happens?

Reveal a model answer

The success transaction finds the old hold ineligible and records refund work without changing the new allocation.

What the answer must demonstrate: Recheck hold eligibility and ownership; compensate without reclaiming another hold’s seats.

Follow-up · Question 6Design a ticket-booking service

Does a waiting room guarantee first-arrival seat allocation?

Reveal a model answer

No. Several admitted clients can race and network timing can reorder their commits. State whether the product promises bounded load or strict fairness.

What the answer must demonstrate: State the chosen admission fairness contract and its concurrency cost.

Foundation · Question 7Design a ticket-booking service

Why partition by show?

Reveal a model answer

Customers compete for seats within one show. Keeping those seats together lets each hold commit atomically, while different shows scale independently.

What the answer must demonstrate: Keep a show’s seat set local while recognizing the single-show contention limit.

Follow-up · Question 8Design a ticket-booking service

Why are three database replicas not a complete durability answer?

Reveal a model answer

The commit and promotion policies determine whether acknowledged holds survive and whether old writers are fenced. Replica count alone says neither.

What the answer must demonstrate: Name acknowledgment, safe promotion and old-writer fencing instead of counting replicas.

Follow-up · Question 9Design a ticket-booking service

How does the design change when the booking is a three-night hotel stay?

Reveal a model answer

Inventory becomes a room-type capacity row for every occupied hotel-local date in [checkIn, checkOut). One transaction checks and reserves all nights. If the middle night is unavailable, reject the entire stay with no partial holds. Confirmation or release changes every nightly allocation and the reservation state atomically.

What the answer must demonstrate: Identify room-type/date inventory, exclusive checkout date, whole-stay atomicity and duplicate-safe hold transitions.

Foundation · Question 1Design a distributed key-value store

How is a version different from a request ID?

Reveal a model answer

The version identifies state to replace; the request ID identifies one attempted change. Their combination supports conflict detection and retry recovery.

What the answer must demonstrate: Distinguish expected state from attempted operation and reject conflicting ID reuse.

Two clients replace version 7. Why can only one win?

Reveal a model answer

The agreed command order applies one check-and-update first, advancing the version. The next command checks the new current state and records conflict.

What the answer must demonstrate: Place version comparison and mutation inside the agreed serialized application step.

A write commits but its reply is lost. What survives?

Reveal a model answer

The value, version and saved result remain recoverable together. Retrying the same scoped request returns that result without another change.

What the answer must demonstrate: Persist value and original outcome together across commit and response loss.

Why might a process labeled leader be unable to serve a strong GET?

Reveal a model answer

It may be isolated while another group has elected a successor. The leader must confirm it still leads and apply the committed commands required for the read.

What the answer must demonstrate: Require current read authority and sufficient applied state after leadership change.

Foundation · Question 5Design a distributed key-value store

Does R+W>N fully specify linearizability?

Reveal a model answer

No. Set overlap alone does not define ordering, election safety, failed writes, application or read authority. Use a complete replication protocol.

What the answer must demonstrate: Explain why quorum overlap needs a complete ordering and election protocol.

Foundation · Question 6Design a distributed key-value store

Why store a deletion marker?

Reveal a model answer

Older values may remain in disk files or replicas. The marker prevents their return until safe reclamation rules allow removal.

What the answer must demonstrate: Prevent resurrection and old-version reuse across deletion and recreation.

Why move retry results with values?

Reveal a model answer

A client may retry through the new owner after losing an old response. Without the saved result, the new owner cannot preserve that retry contract.

What the answer must demonstrate: Move retry history, versions and tombstones with the key’s value.

Will adding partitions accelerate one heavily updated key?

Reveal a model answer

Not its ordered conditional decisions. Partitioning adds parallelism across different keys, not concurrent winners against the same state.

What the answer must demonstrate: Recognize the single-key serialization boundary despite more partitions.

Foundation · Question 1Design a notification service

Why separate intent, delivery and attempt?

Reveal a model answer

One business event can create several channel deliveries, and each delivery may require several calls. Separate identities preserve partial status and safe retries.

What the answer must demonstrate: Keep business intent, channel delivery and transport attempt identities separate.

Foundation · Question 2Design a notification service

Does provider acceptance prove the user read the message?

Reveal a model answer

No. Admission, provider acceptance, device delivery and human reading are distinct facts with channel-specific evidence.

What the answer must demonstrate: Distinguish service acceptance, provider acceptance, device delivery and human reading.

Applied · Question 4Design a notification service

The email call times out after possible acceptance. How do you retry?

Reveal a model answer

Recover the same delivery through supported idempotency or lookup. Without either, keep unknown and apply the category’s explicit duplicate-risk policy.

What the answer must demonstrate: Use original provider identity or lookup and state the unsupported-provider limit.

Foundation · Question 5Design a notification service

Why can in-app notifications have a stronger uniqueness boundary?

Reveal a model answer

The item insertion and completed delivery state can share one local transaction with a unique delivery ID.

What the answer must demonstrate: Identify the local transaction boundary that makes inbox insertion unique.

Applied · Question 6Design a notification service

A campaign fills the queue. How do urgent order updates meet their target?

Reveal a model answer

Reserve provider capacity for transactional work and apply tenant fairness. More workers cannot exceed the same provider quota.

What the answer must demonstrate: Reserve urgency within actual provider quotas and calculate net backlog drain.

Follow-up · Question 7Design a notification service

A rejection arrives for an old push token after refresh. What changes?

Reveal a model answer

Disable only the matching destination version, preserving the newer registration.

What the answer must demonstrate: Guard device invalidation by destination version and freeze uncertain send parameters.

Follow-up · Question 8Design a notification service

How should duplicate or reordered callbacks behave?

Reveal a model answer

Verify and persist them, deduplicate available event identities and apply channel-specific facts without regressing known outcomes.

What the answer must demonstrate: Persist verified provider facts and apply channel-specific idempotent transitions.

Foundation · Question 1Design a payment system and ledger

How does the ledger differ from payment workflow state?

Reveal a model answer

Workflow records which actions are pending or known complete. Immutable journals record financial movements with balanced debit and credit totals per currency.

What the answer must demonstrate: Distinguish mutable workflow state from immutable balanced accounting journals.

Can two different idempotency keys capture one payment twice?

Reveal a model answer

Not if the command also claims a unique per-payment capture slot. Request-key deduplication alone only protects repeats of the same command.

What the answer must demonstrate: Protect both one allowed capture command and one journal per financial operation.

The processor captured but the worker lost its reply. What happens?

Reveal a model answer

Keep the stored operation unknown, recover through its provider identity or lookup, and apply verified evidence locally. Do not issue a new charge to repair missing local status.

What the answer must demonstrate: Recover the original processor effect without issuing another charge to repair local state.

Foundation · Question 4Design a payment system and ledger

Does a balanced journal prove the payment is correct?

Reveal a model answer

No. A wrong amount or merchant can still balance. Validate the operation evidence and compare independent processor facts through reconciliation.

What the answer must demonstrate: Validate merchant, amount, currency and processor evidence in addition to balanced totals.

Two agents each refund 2000 from a 2500 capture. How is excess prevented?

Reveal a model answer

The first atomically reserves 2000 on the capture balance; the second sees only 500 available and cannot call the provider.

What the answer must demonstrate: Reserve shared refund capacity before calls and retain it through unknown outcomes.

Why is provider-event deduplication insufficient for ledger uniqueness?

Reveal a model answer

Different events may describe the same capture. The posting identity must represent the financial operation, not only each callback.

What the answer must demonstrate: Deduplicate financial operations separately from repeated provider event IDs.

Follow-up · Question 7Design a payment system and ledger

Can a reporting replica approve a refund during primary failure?

Reveal a model answer

No. Its balance may omit another refund reservation. The capacity decision needs the current transactional authority.

What the answer must demonstrate: Keep capacity decisions on current authority rather than stale reporting balances.

Follow-up · Question 8Design a payment system and ledger

Why preserve operation IDs in a regional restore?

Reveal a model answer

Restored balances alone do not reveal which external effects already happened. Lost identities can cause replayed commands to charge again.

What the answer must demonstrate: Restore operation identities and reconcile external evidence before resuming commands.

Follow-up · Question 9Design a payment system and ledger

What must change when the prompt asks Alice to transfer existing wallet funds to Bob?

Reveal a model answer

Recover an authorized matching retry by its stable transfer ID before applying new-transfer eligibility checks. For a new positive same-currency transfer, authenticate Alice’s debit authority, lock both accounts in stable order, recheck spendable funds after outgoing holds, and commit the balanced journal, both balances and saved result together. This internal transfer does not need a card-processor call.

What the answer must demonstrate: Separate internal transfers from merchant capture; protect spendable balance and journal atomicity under concurrent spends and retries.

Foundation · Question 1Design a collaborative text editor

Why not save the whole document after each edit?

Reveal a model answer

Two users can replace the same base with different complete copies, causing the later save to erase independent work. Operations preserve the intent the merge algorithm needs.

What the answer must demonstrate: Explain lost independent edits under whole-document replacement.

How do X and Y inserted at position 1 both survive?

Reveal a model answer

A defined tie-break accepts one first and transforms the other position around it. Both clients reconcile pending and accepted edits under the same rules.

What the answer must demonstrate: Trace both client reconciliation and deterministic server transformation without generalizing an insert demo.

Foundation · Question 3Design a collaborative text editor

When can the interface show saved?

Reveal a model answer

After the authoritative operation log durably accepts the edit, not merely after local rendering or gateway receipt.

What the answer must demonstrate: Identify durable acceptance separately from local display and peer delivery.

A17 committed but its reply vanished. How is a duplicate X avoided?

Reveal a model answer

Retry A17 with the same authenticated identity and payload. The owner returns the existing accepted version, and the client removes its pending entry.

What the answer must demonstrate: Reuse operation identity and remove the matching pending edit without applying it twice.

How can an edit disappear between opening and subscribing?

Reveal a model answer

An operation may commit after the initial snapshot/read but before live subscription. Replay or buffer from the last applied version to cover that gap.

What the answer must demonstrate: Close the snapshot-to-subscription gap and repair missing versions before positional application.

Follow-up · Question 6Design a collaborative text editor

Why must storage reject a stale document coordinator?

Reveal a model answer

The old process may resume with open sockets after a replacement takes over. Checking its ownership generation at append prevents a second accepted history.

What the answer must demonstrate: Check ownership generation and expected head at the protected append boundary.

Follow-up · Question 7Design a collaborative text editor

Can a snapshot replace all operation history immediately?

Reveal a model answer

It can reconstruct current text but may not preserve the context needed to transform supported old pending edits. Retention follows the reconnect contract.

What the answer must demonstrate: Retain required transform context or explicitly preserve drafts for resynchronization.

Foundation · Question 8Design a collaborative text editor

Why treat cursors differently from document edits?

Reveal a model answer

Cursors are ephemeral hints that may expire or be coalesced. Accepted edits are durable content requiring replay and uniqueness.

What the answer must demonstrate: Separate ephemeral presence, durable edits and the limits of access revocation.

Why retain metrics, logs and traces separately?

Reveal a model answer

Metrics summarize aggregate behavior, logs describe events and traces link timed work across a request. Each has different identity, storage and query needs.

What the answer must demonstrate: Connect aggregate symptoms, individual events and request-span relationships.

Why is requestId a dangerous metric label?

Reveal a model answer

It can create a new series for every request, exhausting metadata and indexes even with modest sample bytes.

What the answer must demonstrate: Explain series cardinality and place per-request identity in detailed signals.

Why compute rates before summing counters?

Reveal a model answer

Each original series can reset independently. Reset-aware rates preserve those boundaries; summing raw counters first can hide resets and create false activity.

What the answer must demonstrate: Handle each counter reset before combining rates across instances.

Can instance p99 values be averaged?

Reveal a model answer

No. The summaries lose distribution shape and traffic weighting. Merge compatible histograms and compute the combined quantile.

What the answer must demonstrate: Aggregate compatible distributions before estimating a service percentile.

Does durable acceptance mean a record is already searchable?

Reveal a model answer

Not in a buffered design. Accepted input may wait for backend ingestion, so report visibility lag separately.

What the answer must demonstrate: Distinguish pre-acceptance shedding from accepted-data retention and visibility lag.

What should an empty recent window do to an alert?

Reveal a model answer

Apply an explicit stale or missing-data state rather than infer zero errors or automatic recovery.

What the answer must demonstrate: Require fresh evidence and proven duration rather than equating missing data with zero.

A writer returns at normal arrival throughput. When does its outage backlog drain?

Reveal a model answer

It does not. Completion capacity must exceed continuing arrivals, or new work must be reduced.

What the answer must demonstrate: Calculate recovery throughput minus new arrivals and track backlog age.

Does a durable queue alone prevent duplicate stored telemetry?

Reveal a model answer

No. Writers must reuse record IDs, and the backend must handle repeated writes under its documented contract.

What the answer must demonstrate: Name stable record identity and the actual backend output/retry boundary.

Foundation · Question 1Design a distributed job scheduler

How do schedule, occurrence and attempt differ?

Reveal a model answer

The schedule is the rule, an occurrence is one intended resolved run, and attempts retry that same run. This keeps history and business identity stable.

What the answer must demonstrate: Preserve one intended occurrence and its frozen parameters across attempts.

Two scanners see Tuesday due. What prevents two runs?

Reveal a model answer

Both use the unique schedule/instant identity and transactionally create the occurrence, dispatch work and advance next due time.

What the answer must demonstrate: Atomically materialize the due instant, dispatch intent and next schedule time.

Why can an expired worker still be dangerous?

Reveal a model answer

Expiry changes permission but does not stop the process. A paused worker may resume and try to publish after replacement.

What the answer must demonstrate: Enforce the current token at completion because expiry does not stop a process.

Does unique Tuesday occurrence identity prevent Wednesday overlapping?

Reveal a model answer

No. Those are distinct valid occurrences. A schedule-level active-run guard enforces the chosen no-overlap policy across them.

What the answer must demonstrate: Use a separate schedule guard for distinct occurrences and retain physical-execution limits.

Can a unique run row guarantee one email?

Reveal a model answer

No. The email provider performs an external effect. Use a committed stable action identity and the provider’s supported idempotency or reconciliation.

What the answer must demonstrate: Protect external actions independently from local occurrence and result uniqueness.

Foundation · Question 6Design a distributed job scheduler

Why do 50,000 five-second jobs exceed a ten-second start target with 10,000 slots?

Reveal a model answer

They start in five ideal waves at 0, 5, 10, 15 and 20 seconds. The last two waves start late before overhead is counted.

What the answer must demonstrate: Compute execution waves and distinguish start-time targets from completion.

Follow-up · Question 7Design a distributed job scheduler

What must a calendar schedule say about daylight-saving changes?

Reveal a model answer

How nonexistent and repeated local times are handled, using a named zone and resolved execution instant.

What the answer must demonstrate: State time-zone, daylight-saving and recurrence-mode behavior explicitly.

Follow-up · Question 8Design a distributed job scheduler

What does cancellation mean for a running job?

Reveal a model answer

It requests cooperative stopping and prevents later accepted completion/retries once cancellation becomes terminal. External effects already started may still complete.

What the answer must demonstrate: Serialize terminal cancellation with completion and disclose already-started external effects.

Foundation · Question 1Design a distributed message log

Why choose a retained log for search, fraud and analytics?

Reveal a model answer

Each service needs an independent position and replay history. One consumer finishing does not remove the record for the others.

What the answer must demonstrate: Choose independent replay history or per-message task semantics before architecture.

An append reply is lost. Why reuse the sequence?

Reveal a model answer

The append may already be committed. Supported retry identity lets the broker recover the same position instead of treating it as a new append.

What the answer must demonstrate: Reuse supported append identity and retain business identity across producer incarnations.

What if search commits its update and then crashes before its offset?

Reveal a model answer

The replacement replays the event. A unique processed-event identity committed with the update makes that replay harmless.

What the answer must demonstrate: Commit sink deduplication and effect together before advancing broker progress.

Offset 119 finishes before 118. Can the group commit 120?

Reveal a model answer

No. The bookmark must represent a completed prefix; otherwise recovery skips unfinished 118.

What the answer must demonstrate: Advance only through the fully completed partition prefix.

Foundation · Question 5Design a distributed message log

What does a consumer generation protect?

Reveal a model answer

It lets the coordinator reject progress commits from a superseded assignment.

What the answer must demonstrate: Distinguish broker-generation fencing from external sink protection.

Follow-up · Question 6Design a distributed message log

Why can increasing partition count change ordering?

Reveal a model answer

A key may move while older records remain on its previous partition, allowing new events to overtake old ones.

What the answer must demonstrate: Preserve per-key ordering across a route or partition-count change.

Follow-up · Question 7Design a distributed message log

A consumer is eight days behind a seven-day log. What should happen?

Reveal a model answer

Return an explicit retention gap and use an agreed archive or projection rebuild path. Do not silently skip to the newest event.

What the answer must demonstrate: Expose retention gaps and distinguish time history from latest-key compaction.

Does broker transactional processing make an arbitrary payment exactly once?

Reveal a model answer

No. Its guarantee covers only the storage and outputs that participate in its documented transaction protocol; a payment provider needs its own stable identity and uncertain-outcome recovery.

What the answer must demonstrate: Limit exactly-once claims to cooperating storage/effect domains.

What is available stock in this model?

Reveal a model answer

OnHand minus reserved minus allocated. Temporary checkouts and confirmed unshipped orders both claim physical units.

What the answer must demonstrate: Trace onHand, reserved and allocated consistently through shipment.

Two checkouts each request two of three mugs. What prevents overselling?

Reveal a model answer

A short transaction locks or conditionally updates the same stock authority and checks the full line set. After one reserves two, the other sees only one available.

What the answer must demonstrate: Use one authoritative atomic all-line stock decision, not browsing caches.

How does hold expiry race with allocation?

Reveal a model answer

Both guard current hold state and stock counters in one transaction. Allocation first makes expiry a no-op; expiry first makes allocation fail.

What the answer must demonstrate: Guard lifecycle and counters together so retry and expiry cannot release twice.

Capture may have succeeded but timed out. Should allocations expire?

Reveal a model answer

No. Keep the appropriate allocated claim, block shipping and reconcile the original payment attempt. A temporary hold deadline is not a release policy for possibly paid inventory.

What the answer must demonstrate: Keep allocations during unknown capture and reconcile the same financial attempt.

Why freeze the accepted quote?

Reveal a model answer

Prices, tax, discounts and shipping terms may change later. The order must preserve what the customer accepted and the provider should charge.

What the answer must demonstrate: Preserve accepted commercial terms and reject changed-payload key reuse.

How does cancellation race with shipping?

Reveal a model answer

The order transaction commits fulfillment authorization and a unique shipment intent only if allocations and capture are valid. Cancellation must win before that boundary or use an explicit warehouse cancellation/return flow.

What the answer must demonstrate: Order cancellation and fulfillment authorization in one transaction, and create one unique shipment intent.

What changes when lines belong to independent warehouses?

Reveal a model answer

There is no longer one all-line transaction. Persist local step outcomes and compensation while exposing pending state; confirm only after every required allocation exists.

What the answer must demonstrate: State the loss of global atomicity and persist compensating local steps.

Why do warehouse receipts need identities?

Reveal a model answer

Replaying the same receipt must not add physical stock twice. Record the source identity with the stock increase transaction.

What the answer must demonstrate: Deduplicate physical stock receipts and require validated returns before replenishment.

What are the ranks for 920, 900, 900 and 880?

Reveal a model answer

Competition ranks are 1, 2, 2 and 4 because each is one plus the number of strictly higher scores.

What the answer must demonstrate: Define strict-greater competition rank separately from tie-broken display position.

Why is event-ID deduplication insufficient for match awards?

Reveal a model answer

The same match may arrive under a different transport event ID. Store its canonical contribution and source revision so business meaning is applied once.

What the answer must demonstrate: Protect canonical match contribution/revision as well as transport-event identity.

A match contribution changes from 30 to 5. What updates?

Reveal a model answer

Derive delta −25 from the stored contribution and apply it atomically with the newer source revision and player version.

What the answer must demonstrate: Derive correction delta from stored contribution and order projection updates by version.

Why read score and count-above atomically?

Reveal a model answer

Another update between those reads can make the comparison describe different board states. One atomic index operation gives a coherent rank.

What the answer must demonstrate: Read score and population count coherently while disclosing source-to-index lag.

Why send complete totals to the projection?

Reveal a model answer

Repeating a versioned replacement is safely recognizable, whereas an unprotected repeated increment awards points again.

What the answer must demonstrate: Use complete versioned totals and a consistent rebuild/replay boundary.

Why can global top k merge local top k from each player shard?

Reveal a model answer

A player excluded locally has at least k better local players under the same comparator, so cannot be globally top k.

What the answer must demonstrate: Explain local-top-k completeness separately from cross-shard snapshot consistency.

Does waiting one second after cutoff prove an award board is complete?

Reveal a model answer

No. Use the declared late-result policy to establish which accepted inputs count or whether sources have sent all eligible results. Apply those inputs and verify the board before freezing awards.

What the answer must demonstrate: Establish cutoff completeness and retain immutable award/adjudication evidence.

Can a friend board filter the global top 100?

Reveal a model answer

That can miss every eligible friend outside the global list. Rank within the authorized friend population before truncation.

What the answer must demonstrate: Filter the eligible population before ranking or truncating friend results.

Foundation · Question 1Design maps and route planning

What is the first distinction in a maps interview?

Reveal a model answer

Separate drawing a map from computing directions. Tiles are reusable visual objects, while a route is a legal weighted path for particular endpoints and preferences. That distinction explains why CDN bandwidth and search CPU need different capacity estimates. Define departure time and access rules before choosing an engine.

What the answer must demonstrate: Separate drawing a map from computing directions.

Foundation · Question 2Design maps and route planning

Why does A-C-D lose even though A-C is cheapest?

Reveal a model answer

Its full cost is 3 + 8 = 11, while A-B-D costs 4 + 4 = 8. Dijkstra keeps alternative tentative distances, settles C and then B, and improves D before settling it. It does not commit to the cheapest first edge as an entire route. The proof assumes fixed nonnegative costs and a correct legal-state model.

What the answer must demonstrate: Its full cost is 3 + 8 = 11, while A-B-D costs 4 + 4 = 8.

Applied · Question 3Design maps and route planning

What does 2,000 requests/s at 50 ms CPU imply?

Reveal a model answer

It requires 100 CPU-seconds per second, or 100 fully busy cores before headroom. At 60% utilization the rough planning value is 167 cores before redundancy. This is not a deployment guarantee: long routes, memory locality and traffic mix must be benchmarked. Tile delivery remains a separate high-byte workload.

What the answer must demonstrate: It requires 100 CPU-seconds per second, or 100 fully busy cores before headroom.

Applied · Question 4Design maps and route planning

How can an update avoid corrupting an active route?

Reveal a model answer

Build and verify immutable compatible artifacts, then activate their manifest atomically. A request acquires and retains one bundle reference. Cleanup cannot reclaim that bundle until its readers finish. This prevents an active search from combining old shortcuts with incompatible new weights, while permitting later requests to use the new version.

What the answer must demonstrate: Build and verify immutable compatible artifacts, then activate their manifest atomically.

Applied · Question 5Design maps and route planning

B-D closes while the eight-minute route is running. What happens?

Reveal a model answer

The worker validates the expanded road sequence against the closure version at its final boundary. If that version forbids B-D, it recomputes using a compatible model or returns an explicit retry. It cannot release a known-invalid route merely because the original graph was pinned. An incident reported afterward may require a later navigation refresh.

What the answer must demonstrate: The worker validates the expanded road sequence against the closure version at its final boundary.

Follow-up · Question 6Design maps and route planning

Why is choosing the nearest regional border unsafe?

Reveal a model answer

The minimum-cost legal route may cross a farther border, leave and reenter a region, or use a road that looks geometrically indirect. An overlay must represent valid interregional path costs and preserve the search problem. Regional boundaries are deployment boundaries, not road restrictions. Keep full graph replicas when their memory cost is acceptable.

What the answer must demonstrate: The minimum-cost legal route may cross a farther border, leave and reenter a region, or use a road that looks geometrically indirect.

Follow-up · Question 7Design maps and route planning

How would you justify A* or preprocessing?

Reveal a model answer

First measure whether search CPU is the bottleneck. The chosen final design can scale independent local Dijkstra searches with warmed replicas. A* or precomputed shortcuts are further optimizations, tested against that reference. Explain their additional correctness and update assumptions only if the interviewer asks to extend the design.

What the answer must demonstrate: First measure whether search CPU is the bottleneck.

Follow-up · Question 8Design maps and route planning

What do you say when traffic data is unavailable?

Reveal a model answer

Use historical travel weights only if the product permits it, label freshness and preserve known access restrictions. Invalid topology or no legal connection requires failure or a verified base-search fallback, not straight-line directions. Monitor traffic age separately from query success so an available endpoint cannot conceal obsolete estimates.

What the answer must demonstrate: Use historical travel weights only if the product permits it, label freshness and preserve known access restrictions.

Foundation · Question 1Design event-time click analytics

What metric are you actually counting?

Reveal a model answer

Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record. Transport deduplication does not establish that a distinct event is legitimate or billable.

What the answer must demonstrate: Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record.

Foundation · Question 2Design event-time click analytics

What is the smallest complete design?

Reveal a model answer

A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter. Queries read committed counts. Saving both effects together makes a lost worker response recoverable without assuming every delivery occurs once.

What the answer must demonstrate: A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter.

Where does C901 belong when it arrives at 09:02:05?

Reveal a model answer

Its occurrence at 09:00:58 assigns it to the 09:00 minute. If it is within the configured late-update policy, revise that original window from 99 to 100. Otherwise retain it for the correction process. Processing time must not silently answer a different business question.

What the answer must demonstrate: Its occurrence at 09:00:58 assigns it to the 09:00 minute.

The sink commits and the processor crashes before progress is saved. What happens?

Reveal a model answer

Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change. If the previous transaction had aborted, neither identity nor effect would exist and replay would apply it. This contract must be verified for the actual external store.

What the answer must demonstrate: Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change.

How would you handle one ad with 40,000 events/s?

Reveal a model answer

First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total. Stable event identity still prevents duplicate contribution. Splitting every ordinary key adds unnecessary report and state cost.

What the answer must demonstrate: First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total.

Why can worker-local top lists miss an hourly winner?

Reveal a model answer

Red = 6 on two workers totals 12, but each worker can select a different local item at 7. A report must aggregate each ad’s complete total before sorting winners. This design uses a periodic report with stated build time; exact globally simultaneous rankings are a stronger follow-up.

What the answer must demonstrate: Red = 6 on two workers totals 12, but each worker can select a different local item at 7.

Follow-up · Question 7Design event-time click analytics

How long does a ten-minute peak backlog take to recover?

Reveal a model answer

There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds. Preserve admitted work, expose lag and tighten intake before retention is exhausted.

What the answer must demonstrate: There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds.

Follow-up · Question 8Design event-time click analytics

What happens to a very old resend after identity expiry?

Reveal a model answer

A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe. Enforce an admission policy for older work and use controlled reconstruction from retained evidence for historical repair. Do not silently treat forgotten identities as new clicks. Internal recovery obeys the same limit: stop automatic replay beyond retained identity coverage and rebuild affected counters from raw evidence instead of applying old inputs to existing counts.

What the answer must demonstrate: A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe.

Which identifier properties must be clarified first?

Reveal a model answer

Clarify namespace, integer width, allowed gaps, uniqueness and ordering. This design requires distinct integers but permits gaps and non-global issue order, so durable numeric ranges suffice. Record creation, authorization and safe business retries remain separate application responsibilities.

What the answer must demonstrate: Clarify namespace, integer width, allowed gaps, uniqueness and ordering.

How does the single-row baseline avoid duplicates?

Reveal a model answer

A short transaction locks the namespace counter, reserves its next value, advances the counter and commits before returning. Concurrent callers serialize at that row. The durability/failover policy must retain acknowledged advances; unused values may become gaps.

What the answer must demonstrate: A short transaction locks the namespace counter, reserves its next value, advances the counter and commits before returning.

What does reserving 10,000 IDs at a time change?

Reveal a model answer

At ten million IDs/s it reduces authority operations to roughly 1,000 range reservations/s. Each process serves calls with a local synchronized cursor until its range ends. The tradeoff is abandoned space after crashes and loss of global issue order across independent ranges.

What the answer must demonstrate: At ten million IDs/s it reduces authority operations to roughly 1,000 range reservations/s.

Two threads request an ID simultaneously. Why is a range not enough?

Reveal a model answer

Disjoint ranges protect different processes, but both threads could read the same local cursor before either advances it. Use an atomic reservation or lock that checks the end and advances before returning. Batch calls obey the same boundary.

What the answer must demonstrate: Disjoint ranges protect different processes, but both threads could read the same local cursor before either advances it.

How do you recover a generator without persisting every increment?

Reveal a model answer

Discard its entire old remainder and request a fresh disjoint range under a new incarnation. This allows gaps but avoids guessing which values were returned before the crash. A paused old process still holds a different range from its replacement.

What the answer must demonstrate: Discard its entire old remainder and request a fresh disjoint range under a new incarnation.

What makes allocator promotion dangerous?

Reveal a model answer

A replica missing a committed high-water advance can hand out values already reserved elsewhere. Preserve acknowledged range state and request results through the chosen promotion protocol. A stale backup needs equivalent reconciliation or refusal; restarting from its counter silently is unsafe.

What the answer must demonstrate: A replica missing a committed high-water advance can hand out values already reserved elsewhere.

Why may the API return decimal strings?

Reveal a model answer

Some clients cannot exactly represent all supported integers. JavaScript Number loses exactness above 2^53 − 1, so parsing through it can merge distinct IDs. A decimal string and appropriate integer type preserve the value and its namespace.

What the answer must demonstrate: Some clients cannot exactly represent all supported integers.

When would you consider a timestamp-based format?

Reveal a model answer

When compact approximate time ordering is an actual requirement. Timestamp, worker and sequence fields introduce clock rollback, worker reuse and overflow rules that numeric ranges avoid. UUID standards are another alternative when 128-bit storage is acceptable; strict global order still needs coordination.

What the answer must demonstrate: When compact approximate time ordering is an actual requirement.

Foundation · Question 1Design a webhook delivery platform

What does the receiver’s 202 mean?

Reveal a model answer

Under this contract it means the event and processing work were durably accepted. The receiver can respond before completing its business workflow, which preserves low endpoint latency without losing work on process restart. A sender needing proof of completed processing requires a separate status or callback protocol.

What the answer must demonstrate: Under this contract it means the event and processing work were durably accepted.

Foundation · Question 2Design a webhook delivery platform

Which fields change between A1 and A2?

Reveal a model answer

Event E402, delivery D22 and immutable body remain stable. Attempt identity, timestamp, signature and sender lease token change. This lets diagnostics distinguish network attempts while the receiver recognizes the same business event. A new event ID would undermine deduplication.

What the answer must demonstrate: Event E402, delivery D22 and immutable body remain stable.

Can the sender avoid retries after a lost reply?

Reveal a model answer

It cannot infer acceptance from the timeout. Refusing every retry loses events in the history where the request never arrived; retrying can repeat arrival when receipt committed. Receiver inbox identity makes that repetition safe under the agreed contract, or a supported receipt lookup can resolve it.

What the answer must demonstrate: It cannot infer acceptance from the timeout.

A1 reports timeout after A2 succeeds. What changes?

Reveal a model answer

The delivery row changes only if the result update carries the current lease token and expected in-flight state. A stale A1 cannot overwrite A2’s accepted state. Its observation can be retained separately in attempt history. This local fence does not prevent remote duplicate POSTs.

What the answer must demonstrate: The delivery row changes only if the result update carries the current lease token and expected in-flight state.

Version 8 arrives before version 7. May version 7 be ignored?

Reveal a model answer

If events carry complete snapshots, a version guard can safely retain the newer state. If events are dependent deltas, dropping version 7 may lose a necessary effect; detect gaps and replay or fetch authoritative complete state. Waiting for each 202 orders acceptance, not necessarily receiver processing.

What the answer must demonstrate: If events carry complete snapshots, a version guard can safely retain the newer state.

Follow-up · Question 6Design a webhook delivery platform

How does a five-second failing endpoint affect capacity?

Reveal a model answer

At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets. Per-endpoint concurrency and tenant scheduling shares prevent one receiver using the whole pool; global limits protect the fleet. Backoff retains obligations without repeatedly hammering the destination.

What the answer must demonstrate: At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets.

Why validate DNS again on each attempt?

Reveal a model answer

A URL can resolve to a different destination after registration. Validate the actual chosen address and connect to it while checking TLS for the original hostname. Otherwise DNS rebinding or an unchecked second resolution can reach internal resources. Redirects require the same validation or must be disabled.

What the answer must demonstrate: A URL can resolve to a different destination after registration.

Follow-up · Question 8Design a webhook delivery platform

Can an operator safely replay a three-month-old event?

Reveal a model answer

This design retains payloads for seven days, so it rejects unavailable history. A longer archive contract must preserve original bytes and destination ownership, and receiver deduplication or business reconciliation must cover that horizon. Reconstructing the current object is a new event, not faithful replay.

What the answer must demonstrate: This design retains payloads for seven days, so it rejects unavailable history.

Why avoid a remote call for each flag check?

Reveal a model answer

At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency. Local compiled snapshots make lookups cheap and let requests continue briefly during control outages. That benefit requires an explicit stale-use and fallback policy rather than claiming instant global updates.

What the answer must demonstrate: At the assumed billion evaluations/s, network calls become enormous work and add a service dependency to application latency.

Why does expanding ten percent to twenty preserve tenant 54?

Reveal a model answer

Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged. A new random draw each request would not preserve assignment. Tenant-based targeting gives all authorized users of that tenant the same cohort.

What the answer must demonstrate: Its bucket 731 remains below both thresholds 1,000 and 2,000 when seed, targeting key, encoding and hash algorithm stay unchanged.

Why is locking each flag separately insufficient?

Reveal a model answer

A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together. Compile a complete immutable snapshot and retain one reference throughout related evaluations. Atomic pointer replacement changes later requests while old references safely finish.

What the answer must demonstrate: A request can still read F7 after its update and F8 before its dependent update, creating a configuration never validated together.

C18 installs before a slow C17 download. What happens?

Reveal a model answer

The installer validates bytes then compares generations under an atomic update. Since 17 is older than 18, it is ignored. A real rollback also uses a higher generation containing the desired older values. Thus content reversal never requires reversing publication order.

What the answer must demonstrate: The installer validates bytes then compares generations under an atomic update.

Why cannot a successful C17 cache fetch renew freshness?

Reveal a model answer

A cache can return authentic C17 after C18 was published. Periodic synchronization must check the environment publication authority, not merely download old bytes again. The main design tracks synchronization and falls back after a long outage; it does not claim a strict publication-to-disable deadline.

What the answer must demonstrate: A cache can return authentic C17 after C18 was published.

What happens when an instance cannot hear the off switch?

Reveal a model answer

The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization. That is an outage policy, not an instantaneous command or a proved global five-minute disable bound. Sensitive actions need their own current permission check.

What the answer must demonstrate: The application uses its last validated snapshot during a short outage, then returns the documented false fallback after five minutes without successful synchronization.

Is every evaluation an experiment exposure?

Reveal a model answer

No. The decision must influence the experience under the agreed exposure definition. Debug calls, hidden branches and requests that never render can evaluate without exposing a variant. Record actual flag, generation and variant used, and analyze outcome guardrails in addition to adoption.

What the answer must demonstrate: No. The decision must influence the experience under the agreed exposure definition.

How do you validate a multi-language SDK rollout?

Reveal a model answer

Run identical fixed context/seed inputs against expected bucket and rule outputs in every language. Test type and missing-attribute behavior, snapshot installation order, offline expiry and restart. Observe generation spread and fallback rates during a limited rollout. Rule schema/compiler compatibility must be checked before activation.

What the answer must demonstrate: Run identical fixed context/seed inputs against expected bucket and rule outputs in every language.

Foundation · Question 1Design a multitenant SaaS platform

Why is tenant ID in the URL insufficient?

Reveal a model answer

It expresses the requested organization, not membership or permission. Authentication establishes the actor; authorization verifies the active tenant and action. Trusted scope then travels through database queries, caches, jobs and files. An unchecked forwarded tenant header would let a caller choose another customer’s data.

What the answer must demonstrate: It expresses the requested organization, not membership or permission.

Foundation · Question 2Design a multitenant SaaS platform

What does row-level security add?

Reveal a model answer

It lets the database enforce row-access policy as defense in depth. Its assumptions include a constrained application role and initialized transaction-local tenant context. Table owners, bypass privileges or stale pooled-session context can undermine those assumptions. It complements rather than replaces application authorization.

What the answer must demonstrate: It lets the database enforce row-access policy as defense in depth.

Why does request-count limiting miss export overload?

Reveal a model answer

A million-row export can consume roughly 100 CPU-seconds in the example, equivalent to 50,000 ordinary reads. Counting it as one request hides database work and output load. Limit active exports, query duration, scanned/output bytes and connections, with fair tenant scheduling and reserved interactive capacity.

What the answer must demonstrate: A million-row export can consume roughly 100 CPU-seconds in the example, equivalent to 50,000 ordinary reads.

When would you give a tenant dedicated placement?

Reveal a model answer

Use measured sustained demand or explicit region, key, recovery and administration constraints. Ordinary tenants can pool economically in cells. Separate schemas mainly organize namespaces; dedicated compute is needed to isolate certain resource contention. Every placement still requires logical authorization at its access paths.

What the answer must demonstrate: Use measured sustained demand or explicit region, key, recovery and administration constraints.

What if an invoice write overlaps cutover?

Reveal a model answer

The chosen migration pauses tenant writes at a boundary shared by every mutation. It waits for admitted writes to commit or abort before copying the stable data. A writer that completed first is copied with its replay result; one arriving after the freeze is rejected and retries at the new placement. Background writers must follow the same rule.

What the answer must demonstrate: The chosen migration pauses tenant writes at a boundary shared by every mutation.

Follow-up · Question 6Design a multitenant SaaS platform

Activation of C5 times out. May you reopen C2?

Reveal a model answer

No. A timeout does not prove C5 accepted no writes. Keep C2 frozen, inspect durable migration progress and recover or finish the destination transition. Once C5 has new writes, returning to C2 needs a controlled transfer of those changes. The main design accepts a maintenance window rather than simultaneous ambiguous ownership.

What the answer must demonstrate: No. A timeout does not prove C5 accepted no writes.

Does a signed download URL enforce current user identity?

Reveal a model answer

The selected design uses an authenticated gateway that checks current tenant membership and permission on a new download request. An ordinary presigned URL is an alternative bearer capability: possession may allow access until expiry even after revocation. Use that alternative only if its expiry-bound behavior meets the contract.

What the answer must demonstrate: The selected design uses an authenticated gateway that checks current tenant membership and permission on a new download request.

Follow-up · Question 8Design a multitenant SaaS platform

How do you restore one tenant from a pooled backup?

Reveal a model answer

Restore into a separate recovery environment, verify tenant scope and import only intended records through a controlled ownership plan. Include replay outcomes, outbox and placement-control state, then rebuild derived views. Restoring the pooled database in place would overwrite other customers.

What the answer must demonstrate: Restore into a separate recovery environment, verify tenant scope and import only intended records through a controlled ownership plan.

Foundation · Question 1Design a distributed object store

Why separate a key from a version?

Reveal a model answer

The key is the mutable name users request; a version identifies exact immutable bytes. Replacement changes the pointer while an existing reader finishes on its selected version. This prevents mixed downloads.

What the answer must demonstrate: The key is the mutable name users request; a version identifies exact immutable bytes.

What if the process crashes after bytes are written but before publication?

Reveal a model answer

The bytes remain unreferenced temporary work, and the logical object has not changed. A retry can finish the session or cleanup can later reclaim it. Reversing the order could publish a pointer to missing bytes.

What the answer must demonstrate: The bytes remain unreferenced temporary work, and the logical object has not changed.

Foundation · Question 3Design a distributed object store

Why split the 2 GiB example into 64 MiB parts?

Reveal a model answer

It creates 32 independently transferable pieces. The client retries a failed part and can use bounded parallelism. Completion verifies the ordered immutable parts and publishes their manifest without copying the entire video in a transaction.

What the answer must demonstrate: It creates 32 independently transferable pieces. The client retries a failed part and can use bounded parallelism.

Two uploads both expect V4. Who wins?

Reveal a model answer

The metadata authority conditionally updates the key. The first committed replacement wins; the other receives a conflict when V4 is no longer current. This is an explicit lost-update policy rather than accidental last-response-wins behavior.

What the answer must demonstrate: The metadata authority conditionally updates the key.

When may a durable upload be acknowledged?

Reveal a model answer

After all three required independent byte copies are durable and the synchronously replicated metadata publication commits. If a required copy cannot be written, wait or fail completion without publishing it. Merely queueing replication cannot promise survival immediately after success.

What the answer must demonstrate: After all three required independent byte copies are durable and the synchronously replicated metadata publication commits.

V5 is published midway through a V4 range download. What happens?

Reveal a model answer

The existing download continues against the exact V4 manifest and immutable bytes. A new current-key lookup can select V5. Cleanup retains versions needed by active or retained readers.

What the answer must demonstrate: The existing download continues against the exact V4 manifest and immutable bytes.

Follow-up · Question 7Design a distributed object store

Does revoking membership immediately invalidate a signed URL?

Reveal a model answer

Not necessarily. An ordinary signed URL is a bearer grant until its effective expiry. Short expiry limits exposure; an authenticated gateway can enforce current permission on new delivery requests. Neither recalls bytes already received.

What the answer must demonstrate: Not necessarily. An ordinary signed URL is a bearer grant until its effective expiry.

Follow-up · Question 8Design a distributed object store

Why not delete every old upload part after a day?

Reveal a model answer

A completed version may reference that part, and a retained or active reader may still require it. Cleanup must distinguish expired unreferenced uploads from published data. Conservative retention is the simple starting point.

What the answer must demonstrate: A completed version may reference that part, and a retained or active reader may still require it.

Why can the first implementation use ordinary text search?

Reveal a model answer

A relational text index can search five hundred policies. The application retrieves passages, checks current permissions and generates cited answers without a separate vector service. Add semantic search when tests show keyword search misses relevant passages because questions use different wording.

What the answer must demonstrate: Begin with a complete authorized text-search path and justify semantic retrieval using measured candidate recall.

How can a chunking decision make the taxi answer wrong?

Reveal a model answer

If one chunk contains the reimbursement rule and another contains its manager-approval condition, retrieving only the first gives the generator incomplete evidence. Preserve related conditions where possible and include evaluation questions that require them.

What the answer must demonstrate: Connect chunk boundaries to the missing approval condition, rather than treating chunk size as an arbitrary tuning constant.

Where must permission checks occur?

Reveal a model answer

Use tenant filters during search, then check current grants, source version and deletion state before each model receives private passages, including a reranker. Recheck sources before release and when opening citations. Authorize private answer-object access separately; source access does not grant access to another employee’s question.

What the answer must demonstrate: Check private answer-object access separately from evidence access; authorize each private-text recipient and state the in-flight limit.

What happens when version 13 is current but only version 12 is indexed?

Reveal a model answer

The selected contract rejects version 12 for new answers. The assistant may temporarily have insufficient current evidence while indexing catches up. It must not silently represent the older policy as current.

What the answer must demonstrate: Separate the current-version contract from index readiness and describe the resulting temporary evidence gap.

Does a correct citation prove a correct answer?

Reveal a model answer

It proves only that the reference identifies a supplied source. The answer may still reverse its meaning, omit a condition or combine incompatible statements. Validate citation identity and evaluate claim support and task correctness separately.

What the answer must demonstrate: Distinguish reference identity from claim support; preserve the manager-approval qualification in the example.

How does indexing recover after a worker crashes?

Reveal a model answer

The source-version transaction records durable indexing work. A worker retries it with stable document/version/chunk identities, so it can complete missing writes without duplicating passages. The query path still rejects deleted or superseded sources.

What the answer must demonstrate: Use durable follow-up work and stable versioned chunk identities; do not equate source commit with search readiness.

Which capacity estimate matters after search becomes fast?

Reveal a model answer

The model token workload. At 200 requests per second and 4,800 input tokens each, peak input demand is 960,000 tokens per second; 400 output tokens add 80,000 per second. Bound context, output and concurrency rather than relying on request count alone.

What the answer must demonstrate: Calculate input and output token demand and protect interactive capacity from background ingestion.

What should users receive during search, permission and model failures?

Reveal a model answer

If permissions cannot be checked, deny private access. If search is down, report retrieval unavailability rather than claim no evidence exists. If generation is down, currently authorized excerpts can be returned as an explicitly labeled fallback.

What the answer must demonstrate: Distinguish unavailable authority or retrieval from absent evidence, and keep excerpt fallbacks currently authorized.

Foundation · Question 1Design an LLM inference platform

What are prefill and decode, and why measure them separately?

Reveal a model answer

Prefill processes input tokens and builds cached attention state. Decode repeatedly generates the next token using the prefix and cache. A long prefill can delay existing decoders, so first-token latency and gaps between output tokens reveal different scheduling problems.

What the answer must demonstrate: Define both execution phases and connect their interference to first-token and inter-token latency.

Do isolated bounds of 27 prefill workers and 40 decode workers prove forty workers are enough?

Reveal a model answer

No. The isolated measurements each assume a particular workload, while mixed workers share compute, memory bandwidth and capacity between phases. Use those figures as lower bounds, then load-test the combined prompt/output distribution with headroom and latency targets.

What the answer must demonstrate: Treat isolated throughput as lower bounds, then account for mixed execution and live-sequence memory.

Two gateways see the same worker with one free slot. What prevents over-admission?

Reveal a model answer

The worker makes a guarded reservation against its actual sequence and KV capacity before accepting execution. Registry reports guide routing but do not allocate memory. One reservation succeeds; the other must wait within its deadline or use another eligible worker before execution starts.

What the answer must demonstrate: Put the guarded allocation at the worker; explain the simplicity and utilization cost of maximum-length reservation.

Foundation · Question 4Design an LLM inference platform

How do continuous batching and chunked prefill solve different problems?

Reveal a model answer

Continuous batching lets finished sequences leave and new ones enter between iterations, avoiding empty fixed-batch slots. Chunked prefill limits how much long-prompt work runs at once so active decoders get opportunities to produce output.

What the answer must demonstrate: Distinguish dynamic batch membership from prefill scheduling and measure both latency consequences.

Why is it unsafe to free the KV cache as soon as a client disconnects?

Reveal a model answer

A GPU kernel may still be reading those blocks. Immediate reuse by another request can corrupt live execution. Record cancellation, stop future scheduling, wait for in-flight references to finish, and only then release blocks.

What the answer must demonstrate: Preserve in-flight memory references and distinguish recorded cancellation intent from confirmed stopped execution.

Which identity is required for safe prefix reuse?

Reveal a model answer

Exact effective prefix tokens, compatible model weights/tokenizer/template/adapters, and a server-controlled tenant or approved trust scope. Shared blocks remain immutable while live requests reference them. A cache entry is computation state, not an independently authorized answer.

What the answer must demonstrate: Include exact compatible prefix identity and trusted tenant scope, with immutable shared live state.

Follow-up · Question 7Design an LLM inference platform

What survives a worker crash in the selected design?

Reveal a model answer

The generation identity, status and recorded usage survive in durable storage. Live KV tensors and unsaved output events disappear. Mark the dead process’s generations interrupted. A restarted worker has a new incarnation—an identity for that process start—and rejects old assignments; a replacement generation is an explicit new attempt.

What the answer must demonstrate: Separate durable identity from ephemeral state, bind assignments to process incarnations, and make replacement attempts explicit.

Follow-up · Question 8Design an LLM inference platform

How do cumulative usage reports avoid duplicate charges?

Reveal a model answer

Advance the saved count monotonically for one generation: reports of 100 and 120 output tokens yield 120, not 220. Finalize a terminal count so late reports cannot reopen billing. Under this design, work after the last durable report at a crash is an internal unbilled cost.

What the answer must demonstrate: Advance cumulative counts rather than adding them, finalize terminal totals, and disclose unrecorded crash-tail work.

Foundation · Question 1Design durable agent workflows

What makes this more reliable than replaying a chat transcript?

Reveal a model answer

The service stores explicit workflow state, bounded activity results, exact proposals, authenticated approvals and external action identities. A restarted worker resumes from those facts instead of asking the model to reconstruct what happened.

What the answer must demonstrate: The service stores explicit workflow state, bounded activity results, exact proposals, authenticated approvals and external action identities.

Foundation · Question 2Design durable agent workflows

Do 2,083 tasks waiting for approval require 2,083 workers?

Reveal a model answer

No. Waiting tasks are durable records with timers. Workers run ready activities and release capacity while approval is pending. Worker sizing follows active call demand and latency rather than the number of open tasks.

What the answer must demonstrate: No. Waiting tasks are durable records with timers.

Applied · Question 3Design durable agent workflows

P8 was approved, but the vendor changes the price. May the agent submit?

Reveal a model answer

Not under the old approval. Validate the saved quote and exact proposal before admission; changed commercial terms require a new proposal and approval. Generated text saying the new price is acceptable does not authorize it.

What the answer must demonstrate: Not under the old approval. Validate the saved quote and exact proposal before admission; changed commercial terms require a new proposal and approval.

Applied · Question 4Design durable agent workflows

The worker crashes after saving a model result. Should it call the model again?

Reveal a model answer

It should reuse the saved activity result and resume the next state. If the result was never committed, repeating a bounded read or generation may be acceptable; external purchases require a different recovery contract.

What the answer must demonstrate: It should reuse the saved activity result and resume the next state.

Applied · Question 5Design durable agent workflows

The vendor accepted an order but its response was lost. What next?

Reveal a model answer

Recover the same action identity through the provider’s documented idempotent retry or lookup within its supported retention window. Save the confirmed order if found. If it cannot be resolved, report SubmissionUnknown and reconcile; a new key risks another order.

What the answer must demonstrate: Recover the same action identity through the provider’s documented idempotent retry or lookup within its supported retention window.

Follow-up · Question 6Design durable agent workflows

Why is a worker lease insufficient to guarantee one purchase?

Reveal a model answer

An expired worker may already have sent the request, and a remote provider can act after local ownership changes. Leases and task versions coordinate local progress; a stable external action key handles duplicate submission where the provider supports it.

What the answer must demonstrate: An expired worker may already have sent the request, and a remote provider can act after local ownership changes.

Applied · Question 7Design durable agent workflows

Can cancel always promise that no order was created?

Reveal a model answer

Only before submission has been admitted and sent. Afterward the action may already exist remotely. Resolve that action and use the vendor’s cancellation or compensation process under a tracked identity. The UI must distinguish these states.

What the answer must demonstrate: Only before submission has been admitted and sent.

Follow-up · Question 8Design durable agent workflows

A retrieved vendor document tells the model to send credentials elsewhere. What stops it?

Reveal a model answer

Retrieved text is untrusted data. Tool gateways allow only validated operations, enforce current permissions and keep secrets in trusted integrations. The model cannot authorize a destination or grant itself a credential.

What the answer must demonstrate: Retrieved text is untrusted data. Tool gateways allow only validated operations, enforce current permissions and keep secrets in trusted integrations.

Foundation · Question 1Design a recommendation platform

Why is regional popularity a valid starting design?

Reveal a model answer

It provides useful context-sensitive results without historical personalization or a learned model. The complete flow still filters eligibility, limits repeated creators, records response identity and collects actual feedback. It becomes both a comparison baseline and an outage fallback.

What the answer must demonstrate: It provides useful context-sensitive results without historical personalization or a learned model.

Why not score all ten million videos on every request?

Reveal a model answer

At five thousand requests/s that is fifty billion scores/s. At the illustrative 50 microseconds each, it needs about 2.5 million busy cores before other work. Retrieval narrows the pool so richer ranking is affordable.

What the answer must demonstrate: At five thousand requests/s that is fifty billion scores/s.

Foundation · Question 3Design a recommendation platform

What is the difference between candidate retrieval and ranking?

Reveal a model answer

Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity. Ranking spends more information and computation comparing that set. A perfect ranker cannot select a relevant item retrieval never supplied.

What the answer must demonstrate: Retrieval finds a limited set likely to contain useful items using sources such as followed creators, topics, similarity and popularity.

I12 has a high score but U7 completed it. What happens?

Reveal a model answer

Under this product contract it is filtered out. Eligibility and completion rules are explicit constraints, not tiny penalties a sufficiently high engagement score can overcome. Diversity rules then shape the eligible ranked page.

What the answer must demonstrate: Under this product contract it is filtered out.

Why is returning an item not enough to label it as ignored?

Reveal a model answer

The viewer may never have seen it. Record actual visibility separately from the returned list and use labels appropriate to the observation. Unknown exposure is not the same as a negative preference.

What the answer must demonstrate: The viewer may never have seen it.

Follow-up · Question 6Design a recommendation platform

What is future leakage in this design?

Reveal a model answer

Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch. Offline metrics then overstate what serving could have achieved. Use historically appropriate feature values and availability.

What the answer must demonstrate: Training or evaluation uses facts that were not available at the historical decision, such as later popularity or a future watch.

The personalized feature store times out. What should the API return?

Reveal a model answer

Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks. Expose feature age and fallback rate to operations. A relevance outage should not become a reason to serve forbidden content.

What the answer must demonstrate: Use a tested simpler score or contextual popularity within the latency budget, while retaining eligibility checks.

Follow-up · Question 8Design a recommendation platform

What changes when U7 disables personalization?

Reveal a model answer

Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results. Handle stored and derived history under the documented deletion policy and propagation limits.

What the answer must demonstrate: Stop using personal history for candidate retrieval and ranking, invalidate personalized continuation state and serve contextual results.

Why can signaling succeed while the call has no media?

Reveal a model answer

Signaling only exchanges control information. Media still needs compatible settings, a reachable direct or relayed network path and successful encrypted transport. Firewalls or relay exhaustion can block media even when the application socket works.

What the answer must demonstrate: Signaling only exchanges control information. Media still needs compatible settings, a reachable direct or relayed network path and successful encrypted transport.

What are ICE, STUN and TURN responsible for?

Reveal a model answer

ICE gathers and tests candidate paths. STUN helps discover the public address observed outside a local network. TURN relays media when a usable direct path is unavailable. Discovery alone does not supply relay capacity.

What the answer must demonstrate: ICE gathers and tests candidate paths. STUN helps discover the public address observed outside a local network.

What changes when six 1.5 Mbps publishers move from mesh to an SFU?

Reveal a model answer

In mesh each uploads five copies, or 7.5 Mbps. With one stream to the SFU each uploads 1.5 Mbps and SFU ingress is 9 Mbps. Full-quality forwarding to everyone still produces 45 Mbps of server egress.

What the answer must demonstrate: In mesh each uploads five copies, or 7.5 Mbps.

Why would a publisher deliberately upload several qualities?

Reveal a model answer

Simulcast gives the SFU ready-made quality choices for different receivers. It raises publisher encoding work and upload but lets a weak receiver or small tile receive less data without server-side transcoding for every subscription.

What the answer must demonstrate: Simulcast gives the SFU ready-made quality choices for different receivers.

The host removes P6. Which components must react?

Reveal a model answer

The room service saves the authorized decision, and the SFU stops P6’s publication and subscriptions. Reconnect checks current membership. A browser roster update alone does not remove access to forwarded media.

What the answer must demonstrate: The room service saves the authorized decision, and the SFU stops P6’s publication and subscriptions.

What happens after the room’s SFU crashes?

Reveal a model answer

Clients establish fresh transports to a replacement SFU that loads current room permission and subscriptions. There is an interruption; dead in-flight packets and transport state are not transparently restored.

What the answer must demonstrate: Clients establish fresh transports to a replacement SFU that loads current room permission and subscriptions.

Is a call automatically end-to-end encrypted because transport is encrypted?

Reveal a model answer

No. Encrypted browser-to-SFU and SFU-to-browser legs may still let the trusted SFU access media. Encryption that excludes the server changes recording, transcription and moderation capabilities and needs its own key-management design.

What the answer must demonstrate: No. Encrypted browser-to-SFU and SFU-to-browser legs may still let the trusted SFU access media.

Which load test is more useful than opening many signaling sockets?

Reveal a model answer

Realistic packet traffic with subscriptions, loss, bitrate adaptation, relay paths and region capacity reveals the media bottlenecks. Measure audio gaps, video freezes, join delay and reconnect time as well as CPU and egress.

What the answer must demonstrate: Realistic packet traffic with subscriptions, loss, bitrate adaptation, relay paths and region capacity reveals the media bottlenecks.