Recall after learning
Check what you remember.
Use these short cards after studying a lesson. For full interview answers and follow-ups, open the question bank.
System design interview frameworkHow do you justify each component in the diagram?Recall first, then reveal
A responsibility in a user request or a specific capacity/recovery requirement.
Follow one request.
Return to lessonSystem design interview frameworkWhat makes a tradeoff answer complete?Recall first, then reveal
State the choice, the reason, the cost, and how that cost is handled.
Choice → reason → cost → mechanism.
Return to lessonSystem design interview frameworkWhat does a timeout tell the caller?Recall first, then reveal
The result is unknown; the operation may already have committed.
Unknown is not failed.
Return to lessonSystem design interview frameworkWhen should part of an application become a separate service?Recall first, then reveal
When independent scaling, releases or team ownership justify the extra network failures, API compatibility work and data coordination.
Separate for a benefit; explain the added cost.
Return to lessonHTTP APIs and request lifecycleOne request pathRecall first, then reveal
Name lookup → protected connection → routing → identity/permission → data operation → response.
Trace the request.
Return to lessonHTTP APIs and request lifecycleDoes 202 Accepted mean the order is complete?Recall first, then reveal
No. It means the server accepted the work. Return an operation ID or status URL so the client can check progress.
Accepted → check progress → completed.
Return to lessonHTTP APIs and request lifecycleAPI designRecall first, then reveal
Define success, error, retry, size, pagination, and compatibility behavior.
A URL is not the whole contract.
Return to lessonCapacity estimation: throughput, latency, concurrency and storageRate conversionRecall first, then reveal
Daily operations ÷ 86,400 gives average operations per second.
Actions → daily count → seconds.
Return to lessonCapacity estimation: throughput, latency, concurrency and storageAt 2,000 requests/s and 0.05 seconds per request, how many are in progress?Recall first, then reveal
About 100 on average: 2,000 × 0.05. Little’s law requires a stable workload and matching measurement boundaries.
Rate × time = work in progress.
Return to lessonCapacity estimation: throughput, latency, concurrency and storageCache sizeRecall first, then reveal
Distinct hot entries × bytes per entry, then copies and headroom.
Keys, not requests.
Return to lessonDistributed systems: scalability, reliability, availability and efficiencyFive system qualitiesRecall first, then reveal
Scalability: more work; reliability: correct work; availability: usable now; efficiency: resource cost; manageability: safe diagnosis and change.
Ask five different questions.
Return to lessonDistributed systems: scalability, reliability, availability and efficiencyAvailability budgetRecall first, then reveal
Choose a request-based or time-based definition and a window before calculating.
Define the denominator.
Return to lessonDistributed systems: scalability, reliability, availability and efficiencyWhy might adding application servers fail to speed up checkout?Recall first, then reveal
If every server still waits on the same overloaded database, adding servers leaves the bottleneck in place. Distribute or reduce the limiting work.
Find the bottleneck before adding machines.
Return to lessonDatabases, data models, and ACID transactionsWhat comes before SQL versus NoSQL?Recall first, then reveal
The important read/write patterns and the rules that must hold together.
Questions first, product second.
Return to lessonDatabases, data models, and ACID transactionsHow does one purchase explain ACID?Recall first, then reveal
Stock allocation and order creation commit together, preserve rules under the chosen concurrency model, and survive the configured failures.
One purchase, one protected decision.
Return to lessonDatabases, data models, and ACID transactionsDoes flexible schema mean no schema?Recall first, then reveal
Applications still need field meanings, validation, and version compatibility.
Flexible shape still needs shared meaning.
Return to lessonDatabases, data models, and ACID transactionsWhy keep the purchase price in OrderLine?Recall first, then reveal
It records the price committed at purchase time, independently of later catalog changes.
History is a fact, not a live catalog lookup.
Return to lessonDatabase indexes: B-trees, composite keys and query accessHow do you justify the cost of an index?Recall first, then reveal
Database indexes: B-trees, composite keys and query accessFor one customer’s newest orders, what should a composite index put first?Recall first, then reveal
Customer ID narrows the search, followed by the ordering fields and a tie-breaker. Their order must match the query.
Find the customer → order their rows → take the page.
Return to lessonDatabase indexes: B-trees, composite keys and query accessWrite overheadRecall first, then reveal
One row write can update several indexes plus recovery logs.
Every index must be maintained.
Return to lessonStorage engines and data modelsWhat is the difference between model and engine?Recall first, then reveal
The model defines records and access semantics; the engine defines physical pages, files, logs and update behavior.
Meaning above; mechanics below
Return to lessonStorage engines and data modelsWhy does a tombstone exist?Recall first, then reveal
It records deletion so an older value in another file does not become visible again.
A delete must outlive the old copy
Return to lessonStorage engines and data modelsIs a buffered append free of later work?Recall first, then reveal
No. Flushing and compaction convert fast foreground writes into later I/O, CPU and space costs.
Append now, organize later
Return to lessonStorage engines and data modelsDoes WAL imply survival of disk loss?Recall first, then reveal
No. Local recovery logging and off-machine redundancy cover different failures.
Log repairs a crash; copies cover loss
Return to lessonLoad balancing: definition, algorithms and failoverChoose the balancing unitRecall first, then reveal
Connection count, request count, bytes, and CPU work can differ.
Measure the work being balanced.
Return to lessonLoad balancing: definition, algorithms and failoverIf one server fails, what must the remaining servers have?Recall first, then reveal
Spare capacity for redirected traffic, or admission limits that reject excess work. Removing a failed server from routing alone cannot prevent overload.
Detect failure → redirect → check spare capacity.
Return to lessonLoad balancing: definition, algorithms and failoverAffinityRecall first, then reveal
Sending a caller to the same server can reuse its local cache; essential session data must also survive that server failing.
Sticky is not durable.
Return to lessonCaching: cache hits, misses, write policies and invalidationCache questionsRecall first, then reveal
What identifies the entry, which store has the authoritative record, how old may the copy be, and how is it refreshed?
Key → source → allowed age → refresh.
Return to lessonCaching: cache hits, misses, write policies and invalidationInvalidation versus evictionRecall first, then reveal
Invalidation addresses validity; eviction addresses capacity.
Freshness versus space.
Return to lessonCaching: cache hits, misses, write policies and invalidationA reader fetches $20, then a writer saves $25 and clears the cache. What can go wrong?Recall first, then reveal
The delayed reader can refill the cache with $20 after the writer cleared it. The refresh protocol must account for that order of events.
Read old → write new → clear → refill old.
Return to lessonProxies: forward proxy, reverse proxy and API gatewayProxy rolesRecall first, then reveal
Forward proxies serve clients; reverse proxies front origin services.
Who is represented?
Return to lessonProxies: forward proxy, reverse proxy and API gatewayIf a proxy terminates TLS, what must the application verify next?Recall first, then reveal
Protect and authenticate the proxy-to-application connection as needed. Trust forwarded identity headers only from approved proxies that replace untrusted client values.
Check who supplied identity at each hop.
Return to lessonProxies: forward proxy, reverse proxy and API gatewayGateway timeoutRecall first, then reveal
The origin may already have completed the operation.
Recover the result before repeating the action.
Return to lessonData partitioning and shardingWhy choose customer ID as the shard key for order history?Recall first, then reveal
It keeps one customer’s orders together, so one shard can answer the query. Check whether large customers create uneven storage or traffic.
Query together → store together → check imbalance.
Return to lessonData partitioning and shardingShard versus indexRecall first, then reveal
The shard key finds the owner; the index finds records within it.
Which machine, then which record.
Return to lessonData partitioning and shardingOnline movementRecall first, then reveal
Copy, catch up, transfer ownership safely, retire the old copy later.
Copy is not cutover.
Return to lessonConsistent hashing and virtual nodesWhat moves when D40 joins?Recall first, then reveal
Keys in (20,40] change from B50 to D40; P35 moves, while P45 stays with B.
A new token takes only its preceding interval.
Return to lessonConsistent hashing and virtual nodesWhy does P90 belong to A20?Recall first, then reveal
Clockwise search wraps from 99 to 0 and first reaches A20.
The number line closes into a circle.
Return to lessonConsistent hashing and virtual nodesAre virtual nodes extra copies?Recall first, then reveal
No. Several ring positions can belong to one physical machine; redundant copies need separate placement.
Many tokens are not many hosts.
Return to lessonConsistent hashing and virtual nodesDoes balanced key placement solve a viral photo hot key?Recall first, then reveal
No. A single hot key can dominate requests even when key counts are evenly distributed.
Hash spreads keys, not one key’s popularity.
Return to lessonReplication and durabilityWhat must “cart saved” mean?Recall first, then reveal
State where the write must be durably stored before success is returned and which failures it must survive. For example, a safe commit and election protocol can require durable storage on two of three replicas to tolerate one replica loss.
Saved where, before saying saved.
Return to lessonReplication and durabilityWhy might a durable follower show v40?Recall first, then reveal
It may have stored v41 in its log without applying it to queryable state yet.
Received, stored, applied: three milestones.
Return to lessonReplication and durabilityAre replicas a replacement for backups?Recall first, then reveal
No. A bad delete can reach every live copy; retained history is needed to recover the earlier state.
Replicas copy today; backups preserve yesterday.
Return to lessonReplication and durabilityDoes adding read replicas divide the write stream?Recall first, then reveal
No. Copies still process the same writes; sharding divides different data and work.
Copy versus divide.
Return to lessonCAP theorem: consistency, availability, and partition toleranceState CAP and label the triangle.Recall first, then reveal
The CAP theorem states that a distributed read/write system cannot guarantee both consistency (C) and availability (A) when a network partition (P) prevents replicas from communicating. C means linearizability: after a write completes, a later read must return that value or a newer write, as if there were one up-to-date copy. A means every request to a nonfailed participant eventually completes according to the operation’s contract.
Partition present: preserve one history (CP) or complete with weaker consistency (AP). CA excludes the partition case.
Return to lessonCAP theorem: consistency, availability, and partition toleranceBoth replicas are running but cannot exchange messages. Which CAP letter describes this?Recall first, then reveal
P: a network partition. Machines can be alive and serve their local clients while communication between them is lost.
Live participants; unavailable communication.
Return to lessonCAP theorem: consistency, availability, and partition toleranceCan returning “temporarily unavailable” preserve every CAP guarantee?Recall first, then reveal
It can protect an authoritative history, but it sacrifices availability for the refused operation. A quick error is not a successful read of the object.
Fast refusal is still refusal.
Return to lessonCAP theorem: consistency, availability, and partition toleranceDoes a stale seat display necessarily mean the seat can be sold twice?Recall first, then reveal
No. Display reads may be stale while reservations use one atomic authority. CAP read consistency and the no-double-sale business rule are different claims.
Displaying a seat and reserving it can require different consistency guarantees.
Return to lessonConsistency modelsDo CAP consistency, a consistency model and ACID consistency mean the same thing?Recall first, then reveal
No. CAP consistency is linearizability: later reads see a completed write or a newer write. A consistency model defines the allowed visibility and ordering, including weaker models. ACID consistency means correct transactions preserve database and application rules.
CAP: current value. Model: allowed observations. ACID: valid state.
Return to lessonConsistency modelsDoes read-your-writes make all readers current?Recall first, then reveal
No. It preserves the writer’s session promise; client B may still see an older replica unless that reader’s contract requires more.
My save → my view
Return to lessonConsistency modelsWhat distinguishes linearizable from sequential?Recall first, then reveal
Both admit an ordered explanation, but linearizability also preserves real-time order of non-overlapping operations.
Linearizability also respects completed-before-started order.
Return to lessonConsistency modelsDoes eventual mean within a deadline?Recall first, then reveal
No. A deadline needs a separately defined and enforced staleness bound.
Eventually has no stopwatch
Return to lessonConsistency modelsWhat must accompany a dependent reply?Recall first, then reveal
The causal history it depends on must already be visible in the reader’s view.
Cause before consequence
Return to lessonTransaction isolationWhy can two valid snapshot transactions create an invalid roster?Recall first, then reveal
They read the same old set and write different rows, so their combined effect may have no valid serial explanation.
Stable picture ≠ safe decision
Return to lessonTransaction isolationWhat must be retried after a serialization failure?Recall first, then reveal
The entire transaction, including the reads and decision that produced its writes.
Reread, rethink, rewrite
Return to lessonTransaction isolationWhat does a guard row protect?Recall first, then reveal
Only operations that acquire it under the agreed protocol before making protected decisions.
All doors use the same lock
Return to lessonTransaction isolationAre serializable and linearizable identical?Recall first, then reveal
No. Serializability orders transactions by equivalent effects; strict serializability also enforces real-time order.
Serial is order; strict adds time
Return to lessonQuorums, consensus, leases, and fencingWhat does R + W > N establish?Recall first, then reveal
For a fixed set of N replicas, a read contacting R replicas and a write acknowledged by W replicas must share at least one replica when R + W > N. Rules for versions, incomplete writes, and failures are still needed.
Overlap is a building block.
Return to lessonQuorums, consensus, leases, and fencingWhy does W1 need fencing after its lease expires?Recall first, then reveal
It may resume and execute old instructions; the output store must reject its obsolete ownership.
The resource checks the ticket number.
Return to lessonQuorums, consensus, leases, and fencingAre Raft term 4 and E9 epoch 8 the same number?Recall first, then reveal
No. One identifies controller leadership; the other identifies ownership of this job.
Different authorities, different generations.
Return to lessonQuorums, consensus, leases, and fencingDoes a fencing token deduplicate a valid request?Recall first, then reveal
No. A stable operation ID is still needed to identify repeated effects within the same valid ownership epoch.
Who may write is not which attempt this is.
Return to lessonIdempotency, retries, and timeoutsTimeoutRecall first, then reveal
The caller stopped waiting; the durable outcome may already exist.
Unknown is not failed.
Return to lessonIdempotency, retries, and timeoutsThe order was saved, but the reply was lost. What makes the retry safe?Recall first, then reveal
The same authenticated caller, request key and matching request recover the saved order and result. A changed request using that key must be rejected.
Same request → same saved result.
Return to lessonIdempotency, retries, and timeoutsSafe retriesRecall first, then reveal
Safe operation + bounded budget + backoff/jitter + reconciliation.
Retry with a reason and a limit.
Return to lessonMessage queues, event logs, delivery guarantees, and backpressureWhat does the outbox guarantee for P501?Recall first, then reveal
The photo record and intention to publish J501 commit together; a relay can retry publication after failure.
Save the photo record and pending job together.
Return to lessonMessage queues, event logs, delivery guarantees, and backpressureWhy can J501 arrive again after success?Recall first, then reveal
The worker may commit its result but fail before the broker records acknowledgment.
Result committed, receipt on the wire lost.
Return to lessonMessage queues, event logs, delivery guarantees, and backpressureHow long does a 12,000-job backlog take to drain at capacity 400 and arrivals 200 per second?Recall first, then reveal
About 60 seconds under stable-rate assumptions: 12,000 divided by 200 spare jobs/second.
Drain with spare capacity, not total capacity.
Return to lessonMessage queues, event logs, delivery guarantees, and backpressureWhy should a version-2 job not overwrite version 3?Recall first, then reveal
Delivery order and completion order can differ; condition the authoritative update on the intended photo version.
An old job cannot replace a newer photo version.
Return to lessonDistributed transactions and sagasWhat does a remote timeout prove?Recall first, then reveal
Only that the caller did not receive a timely answer; the remote effect may already have succeeded.
Timeout means unknown
Return to lessonDistributed transactions and sagasCan a prepared 2PC participant simply time out and abort?Recall first, then reveal
After voting yes it must learn a safe final decision; unilateral timeout abort can contradict an existing commit decision.
Prepared means promised
Return to lessonDistributed transactions and sagasIs compensation a rollback?Recall first, then reveal
It is a new business action after earlier steps committed, so intermediate observations and irreversible effects remain.
Repair forward, not rewind
Return to lessonDistributed transactions and sagasWhat must survive a worker crash halfway through checkout?Recall first, then reveal
The current workflow step, stable IDs for remote actions, saved pending requests and rules for advancing state safely. A replacement worker can then check uncertain results and continue.
Save progress → retry the same action → check the outcome.
Return to lessonReal-time communication: polling, long polling, SSE, and WebSocketWhy does the receiving client’s five-second poll delay message 501?Recall first, then reveal
The event arrives at 17:00:02, but the next request is at 17:00:05.
Timer decides when to ask.
Return to lessonReal-time communication: polling, long polling, SSE, and WebSocketWhat happens after a long-poll response?Recall first, then reveal
The client immediately issues another request after its last received/applied cursor according to the API contract.
One response per request; then request again.
Return to lessonReal-time communication: polling, long polling, SSE, and WebSocketHow do WebSocket and SSE differ in direction?Recall first, then reveal
WebSocket carries messages both ways; SSE streams server-to-client events while client commands use another request.
Conversation versus broadcast response.
Return to lessonReal-time communication: polling, long polling, SSE, and WebSocketDoes reconnect automatically recover missing messages?Recall first, then reveal
Only if the application stores history, accepts a resume cursor, and handles duplicate replay.
Connection restored ≠ missed messages recovered.
Return to lessonProbabilistic data structuresWhat does a Bloom positive mean?Recall first, then reveal
Every tested position is set; another combination of inserted items may have set them. Verify when correctness requires exact membership.
Positive = possible
Return to lessonProbabilistic data structuresWhat does “no false negatives” assume?Recall first, then reveal
A valid filter covering the queried set, correct insertions and hashing, and no unsafe deletion or state loss.
Trust absence only for the set the filter covers.
Return to lessonProbabilistic data structuresWhich sketch answers which question?Recall first, then reveal
Bloom: membership maybe. HyperLogLog: distinct count. Count-Min: frequency estimate.
Membership → distinct count → frequency.
Return to lessonProbabilistic data structuresCan a Bloom negative prevent two simultaneous inserts?Recall first, then reveal
No. Both workers may see absence; an exact atomic claim or uniqueness rule resolves the race.
A filter is not a lock
Return to lessonKeyword search and vector retrievalThe search index finds a relevant document. May the service return it immediately?Recall first, then reveal
Only after checking that the caller may read the exact content version being returned. Old index permissions may no longer be valid.
Find candidates → check access → return permitted content.
Return to lessonKeyword search and vector retrievalLexical / vectorRecall first, then reveal
Lexical matches terms; vectors match learned similarity; hybrid combines evidence.
Exact words and related meaning.
Return to lessonKeyword search and vector retrievalPrecision / recallRecall first, then reveal
Precision: relevant among returned. Recall: returned among all relevant.
Precision: how useful are the results? Recall: how much was found?
Return to lessonAuthentication, authorization, and tenant isolationAuthentication (AuthN) / authorization (AuthZ)Recall first, then reveal
Authentication, authorization, and tenant isolationAcme and Birch both have invoice I17. What must a cache key include?Recall first, then reveal
The verified tenant ID as well as the invoice ID, with permission scope when users can see different fields. Apply tenant checks to database, search, job and file paths too.
Same record ID can belong to different tenants.
Return to lessonAuthentication, authorization, and tenant isolationEncryptionRecall first, then reveal
Protects bytes and channels; does not decide who may receive decrypted data.
A lock needs a permission rule.
Return to lessonMulti-region architecture and disaster recoveryRPO vs RTORecall first, then reveal
RPO: how far back data may go. RTO: how long recovery may take.
Point = data; time = service.
Return to lessonMulti-region architecture and disaster recoveryWest is ready to take over. Why not redirect traffic immediately?Recall first, then reveal
East may still accept writes. First prevent the old writer from committing, then promote West, route traffic and validate recovery.
Stop old writes → enable new writer → verify.
Return to lessonMulti-region architecture and disaster recoveryBackup vs replicaRecall first, then reveal
A replica follows changes; a protected backup preserves an earlier recovery point.
Copies need history.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryWhat does an SLI measure?Recall first, then reveal
An actual service behavior, such as the fraction of accepted photos ready within 60 seconds. The SLO is the target and evaluation window.
Indicator measures; objective targets.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryHow many bad events does 99.9% permit among one million evaluated photos?Recall first, then reveal
At most 1,000 for that defined metric and window.
One in a thousand is the allowance.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryWhy is P501 a bad completion event despite HTTP success?Recall first, then reveal
Acceptance completed, but the photo waited 95 seconds and became ready at 100 seconds, beyond the 60-second good-event threshold. It consumes error budget; the aggregate SLO depends on all evaluated events.
Upload accepted ≠ photo ready.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryWhen is rollback actually successful?Recall first, then reveal
The known-good configuration is restored, queued work is draining, photos become ready on time, and their contents and access permissions are correct.
Changed back is not yet recovered.
Return to lessonDesign a URL shortenerTwo servers generate q7Lm2Ax9 at the same time. Which step decides who owns it?Recall first, then reveal
Both attempt the unique insert. Only one commits; the other chooses a new candidate. A preliminary “absent” lookup would not prevent the race.
Generate, then claim.
Return to lessonDesign a URL shortenerCreation succeeds, but its response is lost. What must the retry reuse?Recall first, then reveal
Reuse the same request key for that owner; the server returns its saved creation result.
Retry the operation, recover its result.
Return to lessonDesign a URL shortenerAn expired link still occupies storage because cleanup is late. May it redirect?Recall first, then reveal
No. The redirect handler checks expiry during the read; cleanup delay does not extend the link’s lifetime.
Expiry first; reclamation later.
Return to lessonDesign PastebinThe body upload succeeds, but the application crashes before READY. What can readers see, and how does recovery proceed?Recall first, then reveal
Readers see no published paste. A retry verifies the stored body and publishes only if the upload is still active; a canceled attempt stays unavailable.
Bytes first, READY second.
Return to lessonDesign PastebinA cache contains a private paste. Is that enough to return it?Recall first, then reveal
No. The service must still check whether this reader may access the paste.
Stored is not authorized.
Return to lessonDesign PastebinWhen may cleanup delete the body of an abandoned upload?Recall first, then reveal
After saving a cancellation that prevents the uploader from publishing it later.
Cancel publication, then reclaim.
Return to lessonDesign a photo-sharing serviceMaya’s original is stored, but preview generation fails. Is the photo ready, and what should Maya see?Recall first, then reveal
It is not READY. Preserve the original and show processing or a failure; publish only after all required variants exist and their manifest commits.
Original saved does not mean photo ready.
Return to lessonDesign a photo-sharing serviceWhy prepare ordinary authors’ feed entries but fetch celebrity posts when a reader opens the feed?Recall first, then reveal
Preparation saves repeated reads for ordinary audiences. Pulling celebrity posts avoids writing a copy for millions of followers who may never read it.
Spend writes where readers benefit.
Return to lessonDesign a photo-sharing serviceA cached feed entry names a private photo. May the service return it immediately?Recall first, then reveal
No. It identifies a possible feed item; current permission checks decide whether the viewer may receive it.
Candidate first, permission next.
Return to lessonDesign a file synchronization serviceBoth laptops edit revision 12; one commits revision 13. What happens when the other uploads its edit?Recall first, then reveal
The atomic expected-revision check fails. Save the second edit as a conflict copy and record that outcome so retrying does not create another copy.
Compare before replacing.
Return to lessonDesign a file synchronization serviceA device misses a socket notification. How does it find the update later?Recall first, then reveal
Read the committed change log after the device’s saved cursor. The notification only prompts an earlier check.
Hints wake; history recovers.
Return to lessonDesign a file synchronization serviceWhen may cleanup delete a stored chunk?Recall first, then reveal
When no retained revision references it and no still-valid upload can publish a revision that uses it.
No references, no publication right.
Return to lessonDesign a chat messaging serviceThe server saves Nora’s message while Sam’s phone is online and laptop asleep. Which statuses can each device claim?Recall first, then reveal
The server can report accepted after commit. The phone reports delivered after saving the message and cursor; the laptop catches up later. Read is a separate client report, not proof of attention.
Save, deliver, report read.
Return to lessonDesign a chat messaging serviceA send times out. What must the client keep when retrying?Recall first, then reveal
Reuse the same client message ID and original payload so the server can return the first send’s result.
One intent, one identity.
Return to lessonDesign a chat messaging serviceHow does a device recover messages received while it was offline?Recall first, then reveal
Fetch ordered conversation history after its saved cursor, then advance the cursor only through messages it has durably saved without gaps.
Sockets notify; history recovers.
Return to lessonDesign a microblogging serviceWhere is one published post stored, and how does it reach follower feeds?Recall first, then reveal
Save its source record and author-history entry, then create feed references through delivery work that can resume after a crash.
One body, many views.
Return to lessonDesign a microblogging serviceA worker inserts p701 for viewer A, then crashes before viewer B and before saving progress. What should the retry do?Recall first, then reveal
Repeat the follower page. The unique viewer/post key makes A’s insertion harmless, B receives the missing entry, and progress advances only after every page write succeeds.
Write, then checkpoint.
Return to lessonDesign a microblogging serviceA feed still contains a deleted or newly private post. What prevents disclosure?Recall first, then reveal
Before returning it, check the source post’s current visibility and the viewer’s current membership or relationship.
Candidates propose; permissions decide.
Return to lessonDesign a video streaming serviceWhat must happen between an upload and a playable video?Recall first, then reveal
Verify and save the original, save processing work that survives crashes, then publish a complete verified manifest.
Source, job, playable output.
Return to lessonDesign a video streaming serviceOn the example 2 Mb/s link, a four-second 5 Mb/s segment takes ten seconds to download. What changes at 1 Mb/s?Recall first, then reveal
The four-second segment needs 0.5 MB and downloads in about two seconds, so the buffer can recover. The player selects a prepared lower-bitrate segment at the next aligned boundary.
Download faster than playback consumes.
Return to lessonDesign a video streaming serviceHow do we stop an obsolete processing worker from publishing broken playback?Recall first, then reveal
Check that its attempt is still current before accepting its manifest, and keep the immutable files it references protected from deletion.
Guard the pointer and its files.
Return to lessonDesign a typeahead servicePopularity changes while queries read ca. Why build a separate hourly snapshot instead of editing the live shortlists?Recall first, then reveal
A live update can change a term’s score before its prefix shortlist changes. A separate build lets each query use one complete ranking; the cost is delayed popularity updates and memory for both versions.
Prepare once, read consistently.
Return to lessonDesign a typeahead serviceWhy retain terms outside the current top k suggestions?Recall first, then reveal
A lower-ranked term may enter the shortlist when a leading term’s score falls or the term is removed.
Today’s loser may win tomorrow.
Return to lessonDesign a typeahead serviceA response for ca arrives after the user types cap. Does a consistent index snapshot make it safe to display?Recall first, then reveal
No. The snapshot keeps ranking data consistent; the browser must also check the input sequence and discard the old ca response.
Version for data, sequence for typing.
Return to lessonDesign an API rate limiterTwo gateways see two admissions inside the current three-slot window. What prevents both from admitting another request?Recall first, then reveal
Each runs the whole prune, count, compare, append and decision-record operation in one transaction for that quota key. One commits first; the other sees three and denies.
Check and consume together.
Return to lessonDesign an API rate limiterA denied request receives a retry time. Is a slot reserved for that time?Recall first, then reveal
No. The caller may check again then, but other callers may have consumed the available slots.
Retry is a check, not a ticket.
Return to lessonDesign an API rate limiterWhat happens if the limiter cannot determine how much allowance remains?Recall first, then reveal
Strict enforcement pauses new admissions rather than risk exceeding the shared limit. Durable, coordinated usage records make safe recovery possible.
No state, no new strict allowance.
Return to lessonDesign public post searchDoes an inverted-index match mean the post may be returned?Recall first, then reveal
No. The index supplies possible matching IDs; the current source version and permission checks decide what may be shown.
Terms find; source confirms.
Return to lessonDesign public post searchT103 is deleted at version 12, then an old version-11 edit arrives. What must the index retain and check?Recall first, then reveal
Retain delete version 12 and atomically reject version 11 as older. Removing the document without its version would let the delayed edit recreate it.
Remember the delete’s version.
Return to lessonDesign public post searchHow does the next search page continue without reordering the same results?Recall first, then reveal
Use the same unexpired index snapshot and continue after the last result’s creation time and post ID.
Same view, next position.
Return to lessonDesign a web crawlerU17 produces manifest M3 with U18 and U19. The worker inserts U18, then crashes. What survives and what repeats?Recall first, then reveal
M3 survives and is replayed. U18’s unique key returns its existing frontier row; U19 is inserted next. Save manifest progress only after both are durably scheduled.
Save links, then advance progress.
Return to lessonDesign a web crawlerTwo URLs return identical bytes. Which checks detect repeated URLs and repeated bodies?Recall first, then reveal
The canonical URL identifies an address; the content digest identifies its bytes. Keep each URL’s context even when reusing the body.
Same address is not same bytes.
Return to lessonDesign a web crawlerHow can more workers increase crawl speed without overloading one site?Recall first, then reveal
Fetch from independent origins in parallel. Workers targeting the same origin share its timing and concurrency limits; all fetches still pass destination checks.
Parallel across sites, polite within each.
Return to lessonDesign a personalized news feedWhat happens between an author publishing and a reader seeing a ranked page?Recall first, then reveal
Save the post, collect eligible candidates, rank them, then return the page. A saved post alone does not mean a follower has received it.
Store, find, order, return.
Return to lessonDesign a personalized news feedWhy prepare ordinary authors’ entries but fetch celebrity posts when the feed is read?Recall first, then reveal
Preparation saves repeated reads. Pulling celebrity posts avoids writing millions of follower entries that few readers may use.
Push reusable work, pull expensive audiences.
Return to lessonDesign a personalized news feedLeo unfollows Maya while a fanout worker is paused; it later inserts p882. Can the next followed-content feed show it?Recall first, then reveal
No, if unfollow committed before the serving check. Recheck the current relationship and exclude the stale candidate. Ranking or cache membership cannot override that decision.
Candidate does not mean permitted.
Return to lessonDesign nearby place search and friend discoveryMaya is at x=990 and P12 at x=1010. Why must a 50-meter search cross the x=1000 cell boundary?Recall first, then reveal
P12 is twenty meters away. Search every intersecting cell, then test exact distance and rank eligible places.
Nearby can cross a cell boundary: cover, measure, rank.
Return to lessonDesign nearby place search and friend discoveryCan checking current place details find a place missing from the search index?Recall first, then reveal
No. It only checks places already found. Finding an omitted place requires an up-to-date index or all relevant recent changes.
Validation is not discovery.
Return to lessonDesign nearby place search and friend discoveryWhen may the service return a friend’s location?Recall first, then reveal
Only when the viewer currently has permission to see that exact position record.
Nearby is not permission.
Return to lessonDesign a ride-hailing backendD17 accepts R501 and R502 at once. Which records must the transaction protect?Recall first, then reveal
Lock the driver and the requested ride, check the offer, then save both assignments, the replay result and notification work together. The losing transaction sees D17 already assigned.
One driver and one ride: protect both sides in one commit.
Return to lessonDesign a ride-hailing backendWhat ride information survives a lost connection?Recall first, then reveal
The saved trip and acceptance result; reconnecting clients can recover them.
Connection is not ownership.
Return to lessonDesign a ride-hailing backendWhich updates may delivery skip when keeping only the latest?Recall first, then reveal
Intermediate GPS coordinates. Trip state changes must remain saved and recoverable.
Skip old coordinates; recover every trip state change.
Return to lessonDesign a ticket-booking serviceA wants seats 54–56 and B wants 56–57. What happens if A’s hold commits first?Recall first, then reveal
B finds seat 56 owned by A and rolls back the entire request, including seat 57. The cached map cannot override that committed hold.
A shared seat creates a conflict: reserve the whole set or none.
Return to lessonDesign a ticket-booking serviceWhat does a payment timeout mean?Recall first, then reveal
The provider may have charged already. Recover the result of the same payment attempt.
Unknown is not declined.
Return to lessonDesign a ticket-booking servicePayment succeeds after the seats were released. What happens?Recall first, then reveal
Save and retry the refund work; never take seats from a newer hold.
Refund the late charge; keep the newer customer’s seats.
Return to lessonDesign a distributed key-value storeCart-42 changed from version 7 to 8, but the reply was lost. What must survive together?Recall first, then reveal
The value or deletion marker, version, saved request result and applied-log position. Together they let recovery return the original result without changing the cart twice.
A lost reply must not repeat the write: save state and result together.
Return to lessonDesign a distributed key-value storeWhat permits a strong read?Recall first, then reveal
The leader confirms it still leads and has applied the committed commands required for this read.
Reachable is not current.
Return to lessonDesign a distributed key-value storeWhich work can more partitions spread across machines?Recall first, then reveal
Operations on different keys. Competing updates to one key must still follow one order.
Many keys, many owners.
Return to lessonDesign a notification serviceN44 has a saved in-app item, but its email call timed out. What can status honestly report?Recall first, then reveal
The intent is durable, the in-app item is complete and email is unknown until provider evidence resolves it. Neither proves the user read the message.
Saved intent, provider acceptance and user reading need different evidence.
Return to lessonDesign a notification serviceDoes a queued notification still need permission before sending?Recall first, then reveal
Yes. The queue schedules work; the database checks current consent. The provider’s retry rules govern repeated external calls.
Queue work; authorize sends.
Return to lessonDesign a notification serviceWhen does a delivery backlog shrink?Recall first, then reveal
Only when completion capacity exceeds the rate of new deliveries arriving.
New work uses capacity too; only the spare part drains backlog.
Return to lessonDesign a payment system and ledgerC81 may have captured USD 25.00, but its reply is lost. What should the worker recover?Recall first, then reveal
Recover C81 through supported provider lookup or idempotency, then post its verified result. A new charge could bill the customer twice.
A lost reply may hide a successful charge: recover the same operation.
Return to lessonDesign a payment system and ledgerHow much more can a captured payment refund?Recall first, then reveal
Captured amount minus completed refunds minus amounts reserved for unresolved refunds.
Unknown still reserves.
Return to lessonDesign a payment system and ledgerWhy can a balanced journal still be wrong?Recall first, then reveal
Its merchant, amount or processor evidence may be wrong. Validate those facts and reconcile them independently as well as balancing the journal.
Balance and evidence.
Return to lessonDesign a collaborative text editorA and B insert X and Y at position 1 in cat. If X is accepted first, where does Y go?Recall first, then reveal
The OT rule moves Y to position 2, yielding cXYat. Both clients apply the same rule to accepted and pending edits.
An earlier insert shifts later positions: transform the edit, not the whole document.
Return to lessonDesign a collaborative text editorHow can a client recover an accepted edit whose broadcast was lost?Recall first, then reveal
Read the saved operation log; operation IDs let the client recognize edits it has already seen.
A lost broadcast is recoverable because the accepted edit is saved.
Return to lessonDesign a collaborative text editorWhy might a current snapshot be insufficient to merge an old pending edit?Recall first, then reveal
The transformation algorithm may still need the intervening edit history to adjust that pending edit.
Current text is not all history.
Return to lessonDesign a metrics, logging and tracing platformReq81 timed out, then collection failed and the graph became empty. Can the alert declare checkout healthy?Recall first, then reveal
No. The graph lacks evidence; it has not measured zero errors. Mark the result stale or unknown until usable observations return.
Missing is not zero.
Return to lessonDesign a metrics, logging and tracing platformHow should several servers’ p99 latency be combined?Recall first, then reveal
Combine compatible histograms or distributions, then estimate the 99th percentile. Do not average each server’s percentile.
Combine histogram counts, then estimate the percentile.
Return to lessonDesign a metrics, logging and tracing platformWhen does a telemetry backlog shrink after recovery?Recall first, then reveal
Only when processing capacity exceeds the rate of new telemetry arriving.
Backlog needs headroom.
Return to lessonDesign a distributed job schedulerA pauses with token 41; B completes R7 with token 42. What happens when A resumes?Recall first, then reveal
R7 and its frozen parameters stay the same, but the database rejects A’s old token. B’s saved result remains the accepted result.
One intended run may execute twice; only the current attempt can save its result.
Return to lessonDesign a distributed job schedulerHow is a late result from a replaced worker rejected?Recall first, then reveal
The result store checks the ownership token atomically when saving completion and rejects an old token.
The result store checks the token when accepting completion.
Return to lessonDesign a distributed job schedulerDoes a ready queue let more jobs execute at once?Recall first, then reveal
No. It buffers and schedules jobs; workers still need enough execution capacity.
Queue is not compute.
Return to lessonDesign a distributed message logSearch saved M17 at offset 117, then crashed before saving nextOffset 118. What happens on restart?Recall first, then reveal
The group bookmark causes M17 to be read again. The processed-event record saved with the database update prevents applying it twice. An offset identifies position, not completed external work.
Save the effect before progress; make replay harmless with a saved event ID.
Return to lessonDesign a distributed message logWhat prevents a replayed event from updating the database twice?Recall first, then reveal
Save the processed event ID and its database change in one transaction; skip an ID already recorded.
Effect and dedupe together.
Return to lessonDesign a distributed message logCan extra partitions split the work for one strictly ordered key?Recall first, then reveal
No. Keeping that key’s order requires one sequence; extra partitions spread other keys.
More partitions spread different keys; one strictly ordered key still has one sequence.
Return to lessonDesign ecommerce checkout and inventory reservationThree mugs are on hand and a checkout holds two. What changes when those two are allocated, then shipped?Recall first, then reveal
Available stays one: allocation moves two from reserved to allocated; shipment reduces onHand and allocated by two. Available = onHand − reserved − allocated.
Holds and allocations both claim stock; shipment removes the units and their allocation together.
Return to lessonDesign ecommerce checkout and inventory reservationA payment call times out. May checkout release stock or try a new charge?Recall first, then reveal
No. Recover the same payment operation before deciding; the first charge may already have succeeded.
Resolve before reallocating.
Return to lessonDesign ecommerce checkout and inventory reservationWhy is cancellation harder after fulfillment was authorized?Recall first, then reveal
Dispatch may already be underway. Cancellation must coordinate with the warehouse rather than simply releasing the order’s stock.
After dispatch authorization, cancellation must confirm that the warehouse has not shipped.
Return to lessonDesign a leaderboard with exact snapshot ranksScores are 920, 900, 900 and 880. Why does the last player rank 4 rather than 3?Recall first, then reveal
Three players have strictly higher scores, so the last rank is 1 + 3 = 4. The two players at 900 both rank 2; tie display order does not change that.
Count higher-scoring players, including tied players above you.
Return to lessonDesign a leaderboard with exact snapshot ranksA correction lowers P2 from 930 to 905. What prevents an old message from restoring 930?Recall first, then reveal
The stored match revision determines the correction, and the newer player version makes the index keep total 905 even if an older total arrives later.
Newer can be lower.
Return to lessonDesign a leaderboard with exact snapshot ranksWhat should scale before splitting one leaderboard across servers?Recall first, then reveal
Cache shared top lists and place independent boards on different servers.
Scale boards before global order.
Return to lessonDesign maps and route planningq61 finds A-B-D, but B-D closes before the reply. What must happen?Recall first, then reveal
Keep one compatible road-data version for the search, then check known closures before replying. Recompute a legal route or return the stated failure; a cached route needs the same check.
One road model; check closures before replying.
Return to lessonDesign maps and route planningWhich resources grow when route queries, map viewing or retained road versions increase?Recall first, then reveal
Route queries need CPU, map tiles need network bandwidth, and retained road graphs need memory.
Compute, bytes, versions.
Return to lessonDesign maps and route planningDoes the displayed ETA promise when the driver will arrive?Recall first, then reveal
No. It estimates arrival using the chosen road model and the updates received so far.
Model time is estimated time.
Return to lessonDesign event-time click analyticsC901 commits, but its queue acknowledgment is lost. Why does replay leave the count at 100?Recall first, then reveal
The processed-event record and count change committed together. Replay finds that record and skips the increment; if the transaction failed, neither change survived.
Save identity with the count; retries cannot count twice.
Return to lessonDesign event-time click analyticsA click arrives late. Which timestamp chooses its minute?Recall first, then reveal
When it occurred. Arrival and processing timestamps explain the delay.
Count when it happened.
Return to lessonDesign event-time click analyticsAn ad’s counts are split across workers. What must happen before ranking ads?Recall first, then reveal
Combine all partial counts for each ad, then compare the complete totals.
Sum before top.
Return to lessonDesign a distributed unique-ID generatorTwo threads request an ID from A’s range at once. What prevents a duplicate?Recall first, then reveal
The authority gives processes non-overlapping ranges. Inside A, an atomic cursor update or lock gives each thread a different value. Both checks are needed.
Separate ranges between processes; synchronize within each.
Return to lessonDesign a distributed unique-ID generatorA process crashes with unused IDs in its range. May its replacement reuse them?Recall first, then reveal
No. Leave that range unused and reserve a new one, accepting gaps.
Burn uncertain values.
Return to lessonDesign a distributed unique-ID generatorWhat does a unique ID leave for the application to decide?Recall first, then reveal
How to recognize a repeated business request, record creation time, check access and provide gapless numbering if required.
A key is not the whole contract.
Return to lessonDesign a webhook delivery platformNorthHarbor saves E402, but its 202 reply disappears. What changes on retry?Recall first, then reveal
Keep E402 and delivery D22; create a new transport attempt. NorthHarbor finds its saved inbox record, so repeated arrival need not repeat the local business effect.
Lost reply → same event, new attempt.
Return to lessonDesign a webhook delivery platformHow does the receiver stop a repeated event from repeating a local update?Recall first, then reveal
In one transaction, change business state and mark the verified inbox event DONE. A retry then sees that completion.
Effect and DONE commit together.
Return to lessonDesign a webhook delivery platformWhy can the sender not promise that every event finishes its business work?Recall first, then reveal
The receiver may stay unreachable, retention may expire, and completion depends on the receiver and any external provider’s guarantees.
Bound attempts; disclose outcomes.
Return to lessonDesign a feature-flag and configuration platformapp8 installs C18, then its slow C17 download finishes. Which version stays active?Recall first, then reveal
C18 stays active because installation accepts only a newer generation. A rollback also gets a new generation, even when it restores older values. Publication, installation and user exposure are separate events.
Newer generation wins, not the last download.
Return to lessonDesign a feature-flag and configuration platformHow can rollback restore older values without accepting an old download?Recall first, then reveal
Design a feature-flag and configuration platformDoes downloading a cached snapshot prove it is current?Recall first, then reveal
No. Check the current generation with the publication service; cached bytes alone cannot show whether a newer version exists.
Ask which version is current; cached bytes cannot answer.
Return to lessonDesign a multitenant SaaS platformT7 and T8 both have invoice 17. What must every access path check?Recall first, then reveal
Verify the caller may act for the requested tenant. Keep that tenant in database relationships, cache keys, jobs, files and operator checks; the invoice number alone is insufficient.
Same invoice number; different tenant ownership.
Return to lessonDesign a multitenant SaaS platformHow does a maintenance move avoid two writable copies of one tenant?Recall first, then reveal
Stop source writes, finish admitted work, copy and verify the data, then enable only the destination. The move accepts a maintenance pause.
Stop writes → drain → copy and verify → enable destination.
Return to lessonDesign a multitenant SaaS platformWhen should a tenant receive dedicated infrastructure?Recall first, then reveal
When measured resource use disrupts others, or a specific storage location or administrative requirement justifies the cost.
Isolate for a concrete constraint.
Return to lessonDesign a distributed object storeU31’s completion succeeds but its reply is lost. What does retry return?Recall first, then reveal
Return the saved V5 result. Completion first verifies durable bytes, then commits the manifest, current-version pointer and result together. A retry does not publish another version.
Durable bytes → atomic publication → saved reply.
Return to lessonDesign a distributed object storeA new version is published during a download. Which bytes should the download use?Recall first, then reveal
Keep the same immutable version selected at the start, and read every part from it.
One read, one version
Return to lessonDesign a permission-aware RAG knowledge assistantThe taxi answer cites P7 v12 but drops manager approval. Which check catches the problem?Recall first, then reveal
Citation validation checks that P7 v12 was supplied; faithfulness checks whether the answer preserves its approval condition. Authorization separately checks whether the employee may use that passage.
Allowed evidence → valid reference → supported claim.
Return to lessonDesign a permission-aware RAG knowledge assistantWhat must be checked besides whether the answer includes a citation?Recall first, then reveal
Permission to use the passage, whether the citation identifies a supplied source, and whether that source supports the answer.
Allowed, identified, supported.
Return to lessonDesign a permission-aware RAG knowledge assistantSearch is unavailable. May the assistant say no supporting policy exists?Recall first, then reveal
No. It could not search reliably; that does not prove the evidence is absent.
Unavailable is not absent.
Return to lessonDesign an LLM inference platformTwo gateways see one free worker slot. Who decides which generation may start?Recall first, then reveal
The worker reserves its actual KV memory before starting work. Router reports can be stale. Admission also bounds input/output tokens and queue waiting, so request count alone is insufficient.
A capacity report is a hint; the worker reserves memory.
Return to lessonDesign an LLM inference platformThe client cancels while a GPU computation still uses its KV blocks. When can memory be freed?Recall first, then reveal
Stop new work, wait until running computations release the blocks, then free them.
Stop → drain → free.
Return to lessonDesign an LLM inference platformDoes a saved generation record let a new worker resume the exact computation?Recall first, then reveal
No. It does not contain live GPU KV memory or output events that were never saved.
Identity survives; execution may not.
Return to lessonDesign durable agent workflowsThe vendor creates O902 but its reply is lost. What may the replacement worker do?Recall first, then reveal
Recover the saved action using the same provider-supported key or a result lookup. If the provider cannot establish the outcome, keep it unknown for reconciliation rather than send a new purchase.
Lost purchase reply → recover the same action.
Return to lessonDesign durable agent workflowsWhy not use a fresh key after a purchase reply is lost?Recall first, then reveal
The first purchase may already exist. Recover its original action through the provider’s supported retry or lookup; a fresh key can create another order.
Unknown is a state
Return to lessonDesign a recommendation platformI11 is returned at position seven, but U7 closes the page before seeing it. Is that an impression?Recall first, then reveal
No. Record actual visibility separately from returned items. Retrieve a bounded candidate set, filter and rank it, then learn asynchronously from what the viewer actually experienced.
Returned is not seen; observation drives learning.
Return to lessonDesign a recommendation platformDoes including an item in the response count as an impression?Recall first, then reveal
No. Count the agreed event showing that the item was actually visible to the user.
Returned is not seen
Return to lessonDesign a live video-conferencing serviceThe signaling socket connects, but P2 hears nothing. Has the call succeeded?Recall first, then reveal
No. Signaling arranges membership and connection setup; media carries the conversation. Check usable received audio/video and the media network path, not just control connectivity.
Setup succeeds only when usable media arrives.
Return to lessonDesign a live video-conferencing serviceHow does an SFU reduce repeated browser uploads?Recall first, then reveal
Each publisher sends its selected quality streams once. The server forwards the appropriate versions to participants allowed to receive them.
Upload once, forward selectively
Return to lesson