System designby Learnastra

Recall after learning

Check what you remember.

Use these short cards after studying a lesson. For full interview answers and follow-ups, open the question bank.

Practise full answers
System design interview frameworkHow do you justify each component in the diagram?Recall first, then reveal

A responsibility in a user request or a specific capacity/recovery requirement.

Follow one request.

Return to lesson
System design interview frameworkWhat makes a tradeoff answer complete?Recall first, then reveal

State the choice, the reason, the cost, and how that cost is handled.

Choice → reason → cost → mechanism.

Return to lesson
System design interview frameworkWhat does a timeout tell the caller?Recall first, then reveal

The result is unknown; the operation may already have committed.

Unknown is not failed.

Return to lesson
System design interview frameworkWhen should part of an application become a separate service?Recall first, then reveal

When independent scaling, releases or team ownership justify the extra network failures, API compatibility work and data coordination.

Separate for a benefit; explain the added cost.

Return to lesson
HTTP APIs and request lifecycleOne request pathRecall first, then reveal

Name lookup → protected connection → routing → identity/permission → data operation → response.

Trace the request.

Return to lesson
HTTP APIs and request lifecycleDoes 202 Accepted mean the order is complete?Recall first, then reveal

No. It means the server accepted the work. Return an operation ID or status URL so the client can check progress.

Accepted → check progress → completed.

Return to lesson
HTTP APIs and request lifecycleAPI designRecall first, then reveal

Define success, error, retry, size, pagination, and compatibility behavior.

A URL is not the whole contract.

Return to lesson
Capacity estimation: throughput, latency, concurrency and storageRate conversionRecall first, then reveal

Daily operations ÷ 86,400 gives average operations per second.

Actions → daily count → seconds.

Return to lesson
Capacity estimation: throughput, latency, concurrency and storageAt 2,000 requests/s and 0.05 seconds per request, how many are in progress?Recall first, then reveal

About 100 on average: 2,000 × 0.05. Little’s law requires a stable workload and matching measurement boundaries.

Rate × time = work in progress.

Return to lesson
Capacity estimation: throughput, latency, concurrency and storageCache sizeRecall first, then reveal

Distinct hot entries × bytes per entry, then copies and headroom.

Keys, not requests.

Return to lesson
Distributed systems: scalability, reliability, availability and efficiencyFive system qualitiesRecall first, then reveal

Scalability: more work; reliability: correct work; availability: usable now; efficiency: resource cost; manageability: safe diagnosis and change.

Ask five different questions.

Return to lesson
Distributed systems: scalability, reliability, availability and efficiencyAvailability budgetRecall first, then reveal

Choose a request-based or time-based definition and a window before calculating.

Define the denominator.

Return to lesson
Distributed systems: scalability, reliability, availability and efficiencyWhy might adding application servers fail to speed up checkout?Recall first, then reveal

If every server still waits on the same overloaded database, adding servers leaves the bottleneck in place. Distribute or reduce the limiting work.

Find the bottleneck before adding machines.

Return to lesson
Databases, data models, and ACID transactionsWhat comes before SQL versus NoSQL?Recall first, then reveal

The important read/write patterns and the rules that must hold together.

Questions first, product second.

Return to lesson
Databases, data models, and ACID transactionsHow does one purchase explain ACID?Recall first, then reveal

Stock allocation and order creation commit together, preserve rules under the chosen concurrency model, and survive the configured failures.

One purchase, one protected decision.

Return to lesson
Databases, data models, and ACID transactionsDoes flexible schema mean no schema?Recall first, then reveal

Applications still need field meanings, validation, and version compatibility.

Flexible shape still needs shared meaning.

Return to lesson
Databases, data models, and ACID transactionsWhy keep the purchase price in OrderLine?Recall first, then reveal

It records the price committed at purchase time, independently of later catalog changes.

History is a fact, not a live catalog lookup.

Return to lesson
Database indexes: B-trees, composite keys and query accessHow do you justify the cost of an index?Recall first, then reveal

Show the filter, order, and limit that it accelerates.

Query first, index second.

Return to lesson
Database indexes: B-trees, composite keys and query accessFor one customer’s newest orders, what should a composite index put first?Recall first, then reveal

Customer ID narrows the search, followed by the ordering fields and a tie-breaker. Their order must match the query.

Find the customer → order their rows → take the page.

Return to lesson
Database indexes: B-trees, composite keys and query accessWrite overheadRecall first, then reveal

One row write can update several indexes plus recovery logs.

Every index must be maintained.

Return to lesson
Storage engines and data modelsWhat is the difference between model and engine?Recall first, then reveal

The model defines records and access semantics; the engine defines physical pages, files, logs and update behavior.

Meaning above; mechanics below

Return to lesson
Storage engines and data modelsWhy does a tombstone exist?Recall first, then reveal

It records deletion so an older value in another file does not become visible again.

A delete must outlive the old copy

Return to lesson
Storage engines and data modelsIs a buffered append free of later work?Recall first, then reveal

No. Flushing and compaction convert fast foreground writes into later I/O, CPU and space costs.

Append now, organize later

Return to lesson
Storage engines and data modelsDoes WAL imply survival of disk loss?Recall first, then reveal

No. Local recovery logging and off-machine redundancy cover different failures.

Log repairs a crash; copies cover loss

Return to lesson
Load balancing: definition, algorithms and failoverChoose the balancing unitRecall first, then reveal

Connection count, request count, bytes, and CPU work can differ.

Measure the work being balanced.

Return to lesson
Load balancing: definition, algorithms and failoverIf one server fails, what must the remaining servers have?Recall first, then reveal

Spare capacity for redirected traffic, or admission limits that reject excess work. Removing a failed server from routing alone cannot prevent overload.

Detect failure → redirect → check spare capacity.

Return to lesson
Load balancing: definition, algorithms and failoverAffinityRecall first, then reveal

Sending a caller to the same server can reuse its local cache; essential session data must also survive that server failing.

Sticky is not durable.

Return to lesson
Caching: cache hits, misses, write policies and invalidationCache questionsRecall first, then reveal

What identifies the entry, which store has the authoritative record, how old may the copy be, and how is it refreshed?

Key → source → allowed age → refresh.

Return to lesson
Caching: cache hits, misses, write policies and invalidationInvalidation versus evictionRecall first, then reveal

Invalidation addresses validity; eviction addresses capacity.

Freshness versus space.

Return to lesson
Caching: cache hits, misses, write policies and invalidationA reader fetches $20, then a writer saves $25 and clears the cache. What can go wrong?Recall first, then reveal

The delayed reader can refill the cache with $20 after the writer cleared it. The refresh protocol must account for that order of events.

Read old → write new → clear → refill old.

Return to lesson
Proxies: forward proxy, reverse proxy and API gatewayProxy rolesRecall first, then reveal

Forward proxies serve clients; reverse proxies front origin services.

Who is represented?

Return to lesson
Proxies: forward proxy, reverse proxy and API gatewayIf a proxy terminates TLS, what must the application verify next?Recall first, then reveal

Protect and authenticate the proxy-to-application connection as needed. Trust forwarded identity headers only from approved proxies that replace untrusted client values.

Check who supplied identity at each hop.

Return to lesson
Proxies: forward proxy, reverse proxy and API gatewayGateway timeoutRecall first, then reveal

The origin may already have completed the operation.

Recover the result before repeating the action.

Return to lesson
Data partitioning and shardingWhy choose customer ID as the shard key for order history?Recall first, then reveal

It keeps one customer’s orders together, so one shard can answer the query. Check whether large customers create uneven storage or traffic.

Query together → store together → check imbalance.

Return to lesson
Data partitioning and shardingShard versus indexRecall first, then reveal

The shard key finds the owner; the index finds records within it.

Which machine, then which record.

Return to lesson
Data partitioning and shardingOnline movementRecall first, then reveal

Copy, catch up, transfer ownership safely, retire the old copy later.

Copy is not cutover.

Return to lesson
Consistent hashing and virtual nodesWhat moves when D40 joins?Recall first, then reveal

Keys in (20,40] change from B50 to D40; P35 moves, while P45 stays with B.

A new token takes only its preceding interval.

Return to lesson
Consistent hashing and virtual nodesWhy does P90 belong to A20?Recall first, then reveal

Clockwise search wraps from 99 to 0 and first reaches A20.

The number line closes into a circle.

Return to lesson
Consistent hashing and virtual nodesAre virtual nodes extra copies?Recall first, then reveal

No. Several ring positions can belong to one physical machine; redundant copies need separate placement.

Many tokens are not many hosts.

Return to lesson
Consistent hashing and virtual nodesDoes balanced key placement solve a viral photo hot key?Recall first, then reveal

No. A single hot key can dominate requests even when key counts are evenly distributed.

Hash spreads keys, not one key’s popularity.

Return to lesson
Replication and durabilityWhat must “cart saved” mean?Recall first, then reveal

State where the write must be durably stored before success is returned and which failures it must survive. For example, a safe commit and election protocol can require durable storage on two of three replicas to tolerate one replica loss.

Saved where, before saying saved.

Return to lesson
Replication and durabilityWhy might a durable follower show v40?Recall first, then reveal

It may have stored v41 in its log without applying it to queryable state yet.

Received, stored, applied: three milestones.

Return to lesson
Replication and durabilityAre replicas a replacement for backups?Recall first, then reveal

No. A bad delete can reach every live copy; retained history is needed to recover the earlier state.

Replicas copy today; backups preserve yesterday.

Return to lesson
Replication and durabilityDoes adding read replicas divide the write stream?Recall first, then reveal

No. Copies still process the same writes; sharding divides different data and work.

Copy versus divide.

Return to lesson
CAP theorem: consistency, availability, and partition toleranceState CAP and label the triangle.Recall first, then reveal

The CAP theorem states that a distributed read/write system cannot guarantee both consistency (C) and availability (A) when a network partition (P) prevents replicas from communicating. C means linearizability: after a write completes, a later read must return that value or a newer write, as if there were one up-to-date copy. A means every request to a nonfailed participant eventually completes according to the operation’s contract.

Partition present: preserve one history (CP) or complete with weaker consistency (AP). CA excludes the partition case.

Return to lesson
CAP theorem: consistency, availability, and partition toleranceBoth replicas are running but cannot exchange messages. Which CAP letter describes this?Recall first, then reveal

P: a network partition. Machines can be alive and serve their local clients while communication between them is lost.

Live participants; unavailable communication.

Return to lesson
CAP theorem: consistency, availability, and partition toleranceCan returning “temporarily unavailable” preserve every CAP guarantee?Recall first, then reveal

It can protect an authoritative history, but it sacrifices availability for the refused operation. A quick error is not a successful read of the object.

Fast refusal is still refusal.

Return to lesson
CAP theorem: consistency, availability, and partition toleranceDoes a stale seat display necessarily mean the seat can be sold twice?Recall first, then reveal

No. Display reads may be stale while reservations use one atomic authority. CAP read consistency and the no-double-sale business rule are different claims.

Displaying a seat and reserving it can require different consistency guarantees.

Return to lesson
Consistency modelsDo CAP consistency, a consistency model and ACID consistency mean the same thing?Recall first, then reveal

No. CAP consistency is linearizability: later reads see a completed write or a newer write. A consistency model defines the allowed visibility and ordering, including weaker models. ACID consistency means correct transactions preserve database and application rules.

CAP: current value. Model: allowed observations. ACID: valid state.

Return to lesson
Consistency modelsDoes read-your-writes make all readers current?Recall first, then reveal

No. It preserves the writer’s session promise; client B may still see an older replica unless that reader’s contract requires more.

My save → my view

Return to lesson
Consistency modelsWhat distinguishes linearizable from sequential?Recall first, then reveal

Both admit an ordered explanation, but linearizability also preserves real-time order of non-overlapping operations.

Linearizability also respects completed-before-started order.

Return to lesson
Consistency modelsDoes eventual mean within a deadline?Recall first, then reveal

No. A deadline needs a separately defined and enforced staleness bound.

Eventually has no stopwatch

Return to lesson
Consistency modelsWhat must accompany a dependent reply?Recall first, then reveal

The causal history it depends on must already be visible in the reader’s view.

Cause before consequence

Return to lesson
Transaction isolationWhy can two valid snapshot transactions create an invalid roster?Recall first, then reveal

They read the same old set and write different rows, so their combined effect may have no valid serial explanation.

Stable picture ≠ safe decision

Return to lesson
Transaction isolationWhat must be retried after a serialization failure?Recall first, then reveal

The entire transaction, including the reads and decision that produced its writes.

Reread, rethink, rewrite

Return to lesson
Transaction isolationWhat does a guard row protect?Recall first, then reveal

Only operations that acquire it under the agreed protocol before making protected decisions.

All doors use the same lock

Return to lesson
Transaction isolationAre serializable and linearizable identical?Recall first, then reveal

No. Serializability orders transactions by equivalent effects; strict serializability also enforces real-time order.

Serial is order; strict adds time

Return to lesson
Quorums, consensus, leases, and fencingWhat does R + W > N establish?Recall first, then reveal

For a fixed set of N replicas, a read contacting R replicas and a write acknowledged by W replicas must share at least one replica when R + W > N. Rules for versions, incomplete writes, and failures are still needed.

Overlap is a building block.

Return to lesson
Quorums, consensus, leases, and fencingWhy does W1 need fencing after its lease expires?Recall first, then reveal

It may resume and execute old instructions; the output store must reject its obsolete ownership.

The resource checks the ticket number.

Return to lesson
Quorums, consensus, leases, and fencingAre Raft term 4 and E9 epoch 8 the same number?Recall first, then reveal

No. One identifies controller leadership; the other identifies ownership of this job.

Different authorities, different generations.

Return to lesson
Quorums, consensus, leases, and fencingDoes a fencing token deduplicate a valid request?Recall first, then reveal

No. A stable operation ID is still needed to identify repeated effects within the same valid ownership epoch.

Who may write is not which attempt this is.

Return to lesson
Idempotency, retries, and timeoutsTimeoutRecall first, then reveal

The caller stopped waiting; the durable outcome may already exist.

Unknown is not failed.

Return to lesson
Idempotency, retries, and timeoutsThe order was saved, but the reply was lost. What makes the retry safe?Recall first, then reveal

The same authenticated caller, request key and matching request recover the saved order and result. A changed request using that key must be rejected.

Same request → same saved result.

Return to lesson
Idempotency, retries, and timeoutsSafe retriesRecall first, then reveal

Safe operation + bounded budget + backoff/jitter + reconciliation.

Retry with a reason and a limit.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureWhat does the outbox guarantee for P501?Recall first, then reveal

The photo record and intention to publish J501 commit together; a relay can retry publication after failure.

Save the photo record and pending job together.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureWhy can J501 arrive again after success?Recall first, then reveal

The worker may commit its result but fail before the broker records acknowledgment.

Result committed, receipt on the wire lost.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureHow long does a 12,000-job backlog take to drain at capacity 400 and arrivals 200 per second?Recall first, then reveal

About 60 seconds under stable-rate assumptions: 12,000 divided by 200 spare jobs/second.

Drain with spare capacity, not total capacity.

Return to lesson
Message queues, event logs, delivery guarantees, and backpressureWhy should a version-2 job not overwrite version 3?Recall first, then reveal

Delivery order and completion order can differ; condition the authoritative update on the intended photo version.

An old job cannot replace a newer photo version.

Return to lesson
Distributed transactions and sagasWhat does a remote timeout prove?Recall first, then reveal

Only that the caller did not receive a timely answer; the remote effect may already have succeeded.

Timeout means unknown

Return to lesson
Distributed transactions and sagasCan a prepared 2PC participant simply time out and abort?Recall first, then reveal

After voting yes it must learn a safe final decision; unilateral timeout abort can contradict an existing commit decision.

Prepared means promised

Return to lesson
Distributed transactions and sagasIs compensation a rollback?Recall first, then reveal

It is a new business action after earlier steps committed, so intermediate observations and irreversible effects remain.

Repair forward, not rewind

Return to lesson
Distributed transactions and sagasWhat must survive a worker crash halfway through checkout?Recall first, then reveal

The current workflow step, stable IDs for remote actions, saved pending requests and rules for advancing state safely. A replacement worker can then check uncertain results and continue.

Save progress → retry the same action → check the outcome.

Return to lesson
Real-time communication: polling, long polling, SSE, and WebSocketWhy does the receiving client’s five-second poll delay message 501?Recall first, then reveal

The event arrives at 17:00:02, but the next request is at 17:00:05.

Timer decides when to ask.

Return to lesson
Real-time communication: polling, long polling, SSE, and WebSocketWhat happens after a long-poll response?Recall first, then reveal

The client immediately issues another request after its last received/applied cursor according to the API contract.

One response per request; then request again.

Return to lesson
Real-time communication: polling, long polling, SSE, and WebSocketHow do WebSocket and SSE differ in direction?Recall first, then reveal

WebSocket carries messages both ways; SSE streams server-to-client events while client commands use another request.

Conversation versus broadcast response.

Return to lesson
Real-time communication: polling, long polling, SSE, and WebSocketDoes reconnect automatically recover missing messages?Recall first, then reveal

Only if the application stores history, accepts a resume cursor, and handles duplicate replay.

Connection restored ≠ missed messages recovered.

Return to lesson
Probabilistic data structuresWhat does a Bloom positive mean?Recall first, then reveal

Every tested position is set; another combination of inserted items may have set them. Verify when correctness requires exact membership.

Positive = possible

Return to lesson
Probabilistic data structuresWhat does “no false negatives” assume?Recall first, then reveal

A valid filter covering the queried set, correct insertions and hashing, and no unsafe deletion or state loss.

Trust absence only for the set the filter covers.

Return to lesson
Probabilistic data structuresWhich sketch answers which question?Recall first, then reveal

Bloom: membership maybe. HyperLogLog: distinct count. Count-Min: frequency estimate.

Membership → distinct count → frequency.

Return to lesson
Probabilistic data structuresCan a Bloom negative prevent two simultaneous inserts?Recall first, then reveal

No. Both workers may see absence; an exact atomic claim or uniqueness rule resolves the race.

A filter is not a lock

Return to lesson
Keyword search and vector retrievalThe search index finds a relevant document. May the service return it immediately?Recall first, then reveal

Only after checking that the caller may read the exact content version being returned. Old index permissions may no longer be valid.

Find candidates → check access → return permitted content.

Return to lesson
Keyword search and vector retrievalLexical / vectorRecall first, then reveal

Lexical matches terms; vectors match learned similarity; hybrid combines evidence.

Exact words and related meaning.

Return to lesson
Keyword search and vector retrievalPrecision / recallRecall first, then reveal

Precision: relevant among returned. Recall: returned among all relevant.

Precision: how useful are the results? Recall: how much was found?

Return to lesson
Authentication, authorization, and tenant isolationAuthentication (AuthN) / authorization (AuthZ)Recall first, then reveal

Identity first; permission for this action and object second.

Who? May do what?

Return to lesson
Authentication, authorization, and tenant isolationAcme and Birch both have invoice I17. What must a cache key include?Recall first, then reveal

The verified tenant ID as well as the invoice ID, with permission scope when users can see different fields. Apply tenant checks to database, search, job and file paths too.

Same record ID can belong to different tenants.

Return to lesson
Authentication, authorization, and tenant isolationEncryptionRecall first, then reveal

Protects bytes and channels; does not decide who may receive decrypted data.

A lock needs a permission rule.

Return to lesson
Multi-region architecture and disaster recoveryRPO vs RTORecall first, then reveal

RPO: how far back data may go. RTO: how long recovery may take.

Point = data; time = service.

Return to lesson
Multi-region architecture and disaster recoveryWest is ready to take over. Why not redirect traffic immediately?Recall first, then reveal

East may still accept writes. First prevent the old writer from committing, then promote West, route traffic and validate recovery.

Stop old writes → enable new writer → verify.

Return to lesson
Multi-region architecture and disaster recoveryBackup vs replicaRecall first, then reveal

A replica follows changes; a protected backup preserves an earlier recovery point.

Copies need history.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryWhat does an SLI measure?Recall first, then reveal

An actual service behavior, such as the fraction of accepted photos ready within 60 seconds. The SLO is the target and evaluation window.

Indicator measures; objective targets.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryHow many bad events does 99.9% permit among one million evaluated photos?Recall first, then reveal

At most 1,000 for that defined metric and window.

One in a thousand is the allowance.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryWhy is P501 a bad completion event despite HTTP success?Recall first, then reveal

Acceptance completed, but the photo waited 95 seconds and became ready at 100 seconds, beyond the 60-second good-event threshold. It consumes error budget; the aggregate SLO depends on all evaluated events.

Upload accepted ≠ photo ready.

Return to lesson
Production readiness: SLI, SLO, observability, and recoveryWhen is rollback actually successful?Recall first, then reveal

The known-good configuration is restored, queued work is draining, photos become ready on time, and their contents and access permissions are correct.

Changed back is not yet recovered.

Return to lesson
Design a URL shortenerTwo servers generate q7Lm2Ax9 at the same time. Which step decides who owns it?Recall first, then reveal

Both attempt the unique insert. Only one commits; the other chooses a new candidate. A preliminary “absent” lookup would not prevent the race.

Generate, then claim.

Return to lesson
Design a URL shortenerCreation succeeds, but its response is lost. What must the retry reuse?Recall first, then reveal

Reuse the same request key for that owner; the server returns its saved creation result.

Retry the operation, recover its result.

Return to lesson
Design a URL shortenerAn expired link still occupies storage because cleanup is late. May it redirect?Recall first, then reveal

No. The redirect handler checks expiry during the read; cleanup delay does not extend the link’s lifetime.

Expiry first; reclamation later.

Return to lesson
Design PastebinThe body upload succeeds, but the application crashes before READY. What can readers see, and how does recovery proceed?Recall first, then reveal

Readers see no published paste. A retry verifies the stored body and publishes only if the upload is still active; a canceled attempt stays unavailable.

Bytes first, READY second.

Return to lesson
Design PastebinA cache contains a private paste. Is that enough to return it?Recall first, then reveal

No. The service must still check whether this reader may access the paste.

Stored is not authorized.

Return to lesson
Design PastebinWhen may cleanup delete the body of an abandoned upload?Recall first, then reveal

After saving a cancellation that prevents the uploader from publishing it later.

Cancel publication, then reclaim.

Return to lesson
Design a photo-sharing serviceMaya’s original is stored, but preview generation fails. Is the photo ready, and what should Maya see?Recall first, then reveal

It is not READY. Preserve the original and show processing or a failure; publish only after all required variants exist and their manifest commits.

Original saved does not mean photo ready.

Return to lesson
Design a photo-sharing serviceWhy prepare ordinary authors’ feed entries but fetch celebrity posts when a reader opens the feed?Recall first, then reveal

Preparation saves repeated reads for ordinary audiences. Pulling celebrity posts avoids writing a copy for millions of followers who may never read it.

Spend writes where readers benefit.

Return to lesson
Design a photo-sharing serviceA cached feed entry names a private photo. May the service return it immediately?Recall first, then reveal

No. It identifies a possible feed item; current permission checks decide whether the viewer may receive it.

Candidate first, permission next.

Return to lesson
Design a file synchronization serviceBoth laptops edit revision 12; one commits revision 13. What happens when the other uploads its edit?Recall first, then reveal

The atomic expected-revision check fails. Save the second edit as a conflict copy and record that outcome so retrying does not create another copy.

Compare before replacing.

Return to lesson
Design a file synchronization serviceA device misses a socket notification. How does it find the update later?Recall first, then reveal

Read the committed change log after the device’s saved cursor. The notification only prompts an earlier check.

Hints wake; history recovers.

Return to lesson
Design a file synchronization serviceWhen may cleanup delete a stored chunk?Recall first, then reveal

When no retained revision references it and no still-valid upload can publish a revision that uses it.

No references, no publication right.

Return to lesson
Design a chat messaging serviceThe server saves Nora’s message while Sam’s phone is online and laptop asleep. Which statuses can each device claim?Recall first, then reveal

The server can report accepted after commit. The phone reports delivered after saving the message and cursor; the laptop catches up later. Read is a separate client report, not proof of attention.

Save, deliver, report read.

Return to lesson
Design a chat messaging serviceA send times out. What must the client keep when retrying?Recall first, then reveal

Reuse the same client message ID and original payload so the server can return the first send’s result.

One intent, one identity.

Return to lesson
Design a chat messaging serviceHow does a device recover messages received while it was offline?Recall first, then reveal

Fetch ordered conversation history after its saved cursor, then advance the cursor only through messages it has durably saved without gaps.

Sockets notify; history recovers.

Return to lesson
Design a microblogging serviceWhere is one published post stored, and how does it reach follower feeds?Recall first, then reveal

Save its source record and author-history entry, then create feed references through delivery work that can resume after a crash.

One body, many views.

Return to lesson
Design a microblogging serviceA worker inserts p701 for viewer A, then crashes before viewer B and before saving progress. What should the retry do?Recall first, then reveal

Repeat the follower page. The unique viewer/post key makes A’s insertion harmless, B receives the missing entry, and progress advances only after every page write succeeds.

Write, then checkpoint.

Return to lesson
Design a microblogging serviceA feed still contains a deleted or newly private post. What prevents disclosure?Recall first, then reveal

Before returning it, check the source post’s current visibility and the viewer’s current membership or relationship.

Candidates propose; permissions decide.

Return to lesson
Design a video streaming serviceWhat must happen between an upload and a playable video?Recall first, then reveal

Verify and save the original, save processing work that survives crashes, then publish a complete verified manifest.

Source, job, playable output.

Return to lesson
Design a video streaming serviceOn the example 2 Mb/s link, a four-second 5 Mb/s segment takes ten seconds to download. What changes at 1 Mb/s?Recall first, then reveal

The four-second segment needs 0.5 MB and downloads in about two seconds, so the buffer can recover. The player selects a prepared lower-bitrate segment at the next aligned boundary.

Download faster than playback consumes.

Return to lesson
Design a video streaming serviceHow do we stop an obsolete processing worker from publishing broken playback?Recall first, then reveal

Check that its attempt is still current before accepting its manifest, and keep the immutable files it references protected from deletion.

Guard the pointer and its files.

Return to lesson
Design a typeahead servicePopularity changes while queries read ca. Why build a separate hourly snapshot instead of editing the live shortlists?Recall first, then reveal

A live update can change a term’s score before its prefix shortlist changes. A separate build lets each query use one complete ranking; the cost is delayed popularity updates and memory for both versions.

Prepare once, read consistently.

Return to lesson
Design a typeahead serviceWhy retain terms outside the current top k suggestions?Recall first, then reveal

A lower-ranked term may enter the shortlist when a leading term’s score falls or the term is removed.

Today’s loser may win tomorrow.

Return to lesson
Design a typeahead serviceA response for ca arrives after the user types cap. Does a consistent index snapshot make it safe to display?Recall first, then reveal

No. The snapshot keeps ranking data consistent; the browser must also check the input sequence and discard the old ca response.

Version for data, sequence for typing.

Return to lesson
Design an API rate limiterTwo gateways see two admissions inside the current three-slot window. What prevents both from admitting another request?Recall first, then reveal

Each runs the whole prune, count, compare, append and decision-record operation in one transaction for that quota key. One commits first; the other sees three and denies.

Check and consume together.

Return to lesson
Design an API rate limiterA denied request receives a retry time. Is a slot reserved for that time?Recall first, then reveal

No. The caller may check again then, but other callers may have consumed the available slots.

Retry is a check, not a ticket.

Return to lesson
Design an API rate limiterWhat happens if the limiter cannot determine how much allowance remains?Recall first, then reveal

Strict enforcement pauses new admissions rather than risk exceeding the shared limit. Durable, coordinated usage records make safe recovery possible.

No state, no new strict allowance.

Return to lesson
Design public post searchDoes an inverted-index match mean the post may be returned?Recall first, then reveal

No. The index supplies possible matching IDs; the current source version and permission checks decide what may be shown.

Terms find; source confirms.

Return to lesson
Design public post searchT103 is deleted at version 12, then an old version-11 edit arrives. What must the index retain and check?Recall first, then reveal

Retain delete version 12 and atomically reject version 11 as older. Removing the document without its version would let the delayed edit recreate it.

Remember the delete’s version.

Return to lesson
Design public post searchHow does the next search page continue without reordering the same results?Recall first, then reveal

Use the same unexpired index snapshot and continue after the last result’s creation time and post ID.

Same view, next position.

Return to lesson
Design a web crawlerU17 produces manifest M3 with U18 and U19. The worker inserts U18, then crashes. What survives and what repeats?Recall first, then reveal

M3 survives and is replayed. U18’s unique key returns its existing frontier row; U19 is inserted next. Save manifest progress only after both are durably scheduled.

Save links, then advance progress.

Return to lesson
Design a web crawlerTwo URLs return identical bytes. Which checks detect repeated URLs and repeated bodies?Recall first, then reveal

The canonical URL identifies an address; the content digest identifies its bytes. Keep each URL’s context even when reusing the body.

Same address is not same bytes.

Return to lesson
Design a web crawlerHow can more workers increase crawl speed without overloading one site?Recall first, then reveal

Fetch from independent origins in parallel. Workers targeting the same origin share its timing and concurrency limits; all fetches still pass destination checks.

Parallel across sites, polite within each.

Return to lesson
Design a personalized news feedWhat happens between an author publishing and a reader seeing a ranked page?Recall first, then reveal

Save the post, collect eligible candidates, rank them, then return the page. A saved post alone does not mean a follower has received it.

Store, find, order, return.

Return to lesson
Design a personalized news feedWhy prepare ordinary authors’ entries but fetch celebrity posts when the feed is read?Recall first, then reveal

Preparation saves repeated reads. Pulling celebrity posts avoids writing millions of follower entries that few readers may use.

Push reusable work, pull expensive audiences.

Return to lesson
Design a personalized news feedLeo unfollows Maya while a fanout worker is paused; it later inserts p882. Can the next followed-content feed show it?Recall first, then reveal

No, if unfollow committed before the serving check. Recheck the current relationship and exclude the stale candidate. Ranking or cache membership cannot override that decision.

Candidate does not mean permitted.

Return to lesson
Design nearby place search and friend discoveryMaya is at x=990 and P12 at x=1010. Why must a 50-meter search cross the x=1000 cell boundary?Recall first, then reveal

P12 is twenty meters away. Search every intersecting cell, then test exact distance and rank eligible places.

Nearby can cross a cell boundary: cover, measure, rank.

Return to lesson
Design nearby place search and friend discoveryCan checking current place details find a place missing from the search index?Recall first, then reveal

No. It only checks places already found. Finding an omitted place requires an up-to-date index or all relevant recent changes.

Validation is not discovery.

Return to lesson
Design nearby place search and friend discoveryWhen may the service return a friend’s location?Recall first, then reveal

Only when the viewer currently has permission to see that exact position record.

Nearby is not permission.

Return to lesson
Design a ride-hailing backendD17 accepts R501 and R502 at once. Which records must the transaction protect?Recall first, then reveal

Lock the driver and the requested ride, check the offer, then save both assignments, the replay result and notification work together. The losing transaction sees D17 already assigned.

One driver and one ride: protect both sides in one commit.

Return to lesson
Design a ride-hailing backendWhat ride information survives a lost connection?Recall first, then reveal

The saved trip and acceptance result; reconnecting clients can recover them.

Connection is not ownership.

Return to lesson
Design a ride-hailing backendWhich updates may delivery skip when keeping only the latest?Recall first, then reveal

Intermediate GPS coordinates. Trip state changes must remain saved and recoverable.

Skip old coordinates; recover every trip state change.

Return to lesson
Design a ticket-booking serviceA wants seats 54–56 and B wants 56–57. What happens if A’s hold commits first?Recall first, then reveal

B finds seat 56 owned by A and rolls back the entire request, including seat 57. The cached map cannot override that committed hold.

A shared seat creates a conflict: reserve the whole set or none.

Return to lesson
Design a ticket-booking serviceWhat does a payment timeout mean?Recall first, then reveal

The provider may have charged already. Recover the result of the same payment attempt.

Unknown is not declined.

Return to lesson
Design a ticket-booking servicePayment succeeds after the seats were released. What happens?Recall first, then reveal

Save and retry the refund work; never take seats from a newer hold.

Refund the late charge; keep the newer customer’s seats.

Return to lesson
Design a distributed key-value storeCart-42 changed from version 7 to 8, but the reply was lost. What must survive together?Recall first, then reveal

The value or deletion marker, version, saved request result and applied-log position. Together they let recovery return the original result without changing the cart twice.

A lost reply must not repeat the write: save state and result together.

Return to lesson
Design a distributed key-value storeWhat permits a strong read?Recall first, then reveal

The leader confirms it still leads and has applied the committed commands required for this read.

Reachable is not current.

Return to lesson
Design a distributed key-value storeWhich work can more partitions spread across machines?Recall first, then reveal

Operations on different keys. Competing updates to one key must still follow one order.

Many keys, many owners.

Return to lesson
Design a notification serviceN44 has a saved in-app item, but its email call timed out. What can status honestly report?Recall first, then reveal

The intent is durable, the in-app item is complete and email is unknown until provider evidence resolves it. Neither proves the user read the message.

Saved intent, provider acceptance and user reading need different evidence.

Return to lesson
Design a notification serviceDoes a queued notification still need permission before sending?Recall first, then reveal

Yes. The queue schedules work; the database checks current consent. The provider’s retry rules govern repeated external calls.

Queue work; authorize sends.

Return to lesson
Design a notification serviceWhen does a delivery backlog shrink?Recall first, then reveal

Only when completion capacity exceeds the rate of new deliveries arriving.

New work uses capacity too; only the spare part drains backlog.

Return to lesson
Design a payment system and ledgerC81 may have captured USD 25.00, but its reply is lost. What should the worker recover?Recall first, then reveal

Recover C81 through supported provider lookup or idempotency, then post its verified result. A new charge could bill the customer twice.

A lost reply may hide a successful charge: recover the same operation.

Return to lesson
Design a payment system and ledgerHow much more can a captured payment refund?Recall first, then reveal

Captured amount minus completed refunds minus amounts reserved for unresolved refunds.

Unknown still reserves.

Return to lesson
Design a payment system and ledgerWhy can a balanced journal still be wrong?Recall first, then reveal

Its merchant, amount or processor evidence may be wrong. Validate those facts and reconcile them independently as well as balancing the journal.

Balance and evidence.

Return to lesson
Design a collaborative text editorA and B insert X and Y at position 1 in cat. If X is accepted first, where does Y go?Recall first, then reveal

The OT rule moves Y to position 2, yielding cXYat. Both clients apply the same rule to accepted and pending edits.

An earlier insert shifts later positions: transform the edit, not the whole document.

Return to lesson
Design a collaborative text editorHow can a client recover an accepted edit whose broadcast was lost?Recall first, then reveal

Read the saved operation log; operation IDs let the client recognize edits it has already seen.

A lost broadcast is recoverable because the accepted edit is saved.

Return to lesson
Design a collaborative text editorWhy might a current snapshot be insufficient to merge an old pending edit?Recall first, then reveal

The transformation algorithm may still need the intervening edit history to adjust that pending edit.

Current text is not all history.

Return to lesson
Design a metrics, logging and tracing platformReq81 timed out, then collection failed and the graph became empty. Can the alert declare checkout healthy?Recall first, then reveal

No. The graph lacks evidence; it has not measured zero errors. Mark the result stale or unknown until usable observations return.

Missing is not zero.

Return to lesson
Design a metrics, logging and tracing platformHow should several servers’ p99 latency be combined?Recall first, then reveal

Combine compatible histograms or distributions, then estimate the 99th percentile. Do not average each server’s percentile.

Combine histogram counts, then estimate the percentile.

Return to lesson
Design a metrics, logging and tracing platformWhen does a telemetry backlog shrink after recovery?Recall first, then reveal

Only when processing capacity exceeds the rate of new telemetry arriving.

Backlog needs headroom.

Return to lesson
Design a distributed job schedulerA pauses with token 41; B completes R7 with token 42. What happens when A resumes?Recall first, then reveal

R7 and its frozen parameters stay the same, but the database rejects A’s old token. B’s saved result remains the accepted result.

One intended run may execute twice; only the current attempt can save its result.

Return to lesson
Design a distributed job schedulerHow is a late result from a replaced worker rejected?Recall first, then reveal

The result store checks the ownership token atomically when saving completion and rejects an old token.

The result store checks the token when accepting completion.

Return to lesson
Design a distributed job schedulerDoes a ready queue let more jobs execute at once?Recall first, then reveal

No. It buffers and schedules jobs; workers still need enough execution capacity.

Queue is not compute.

Return to lesson
Design a distributed message logSearch saved M17 at offset 117, then crashed before saving nextOffset 118. What happens on restart?Recall first, then reveal

The group bookmark causes M17 to be read again. The processed-event record saved with the database update prevents applying it twice. An offset identifies position, not completed external work.

Save the effect before progress; make replay harmless with a saved event ID.

Return to lesson
Design a distributed message logWhat prevents a replayed event from updating the database twice?Recall first, then reveal

Save the processed event ID and its database change in one transaction; skip an ID already recorded.

Effect and dedupe together.

Return to lesson
Design a distributed message logCan extra partitions split the work for one strictly ordered key?Recall first, then reveal

No. Keeping that key’s order requires one sequence; extra partitions spread other keys.

More partitions spread different keys; one strictly ordered key still has one sequence.

Return to lesson
Design ecommerce checkout and inventory reservationThree mugs are on hand and a checkout holds two. What changes when those two are allocated, then shipped?Recall first, then reveal

Available stays one: allocation moves two from reserved to allocated; shipment reduces onHand and allocated by two. Available = onHand − reserved − allocated.

Holds and allocations both claim stock; shipment removes the units and their allocation together.

Return to lesson
Design ecommerce checkout and inventory reservationA payment call times out. May checkout release stock or try a new charge?Recall first, then reveal

No. Recover the same payment operation before deciding; the first charge may already have succeeded.

Resolve before reallocating.

Return to lesson
Design ecommerce checkout and inventory reservationWhy is cancellation harder after fulfillment was authorized?Recall first, then reveal

Dispatch may already be underway. Cancellation must coordinate with the warehouse rather than simply releasing the order’s stock.

After dispatch authorization, cancellation must confirm that the warehouse has not shipped.

Return to lesson
Design a leaderboard with exact snapshot ranksScores are 920, 900, 900 and 880. Why does the last player rank 4 rather than 3?Recall first, then reveal

Three players have strictly higher scores, so the last rank is 1 + 3 = 4. The two players at 900 both rank 2; tie display order does not change that.

Count higher-scoring players, including tied players above you.

Return to lesson
Design a leaderboard with exact snapshot ranksA correction lowers P2 from 930 to 905. What prevents an old message from restoring 930?Recall first, then reveal

The stored match revision determines the correction, and the newer player version makes the index keep total 905 even if an older total arrives later.

Newer can be lower.

Return to lesson
Design a leaderboard with exact snapshot ranksWhat should scale before splitting one leaderboard across servers?Recall first, then reveal

Cache shared top lists and place independent boards on different servers.

Scale boards before global order.

Return to lesson
Design maps and route planningq61 finds A-B-D, but B-D closes before the reply. What must happen?Recall first, then reveal

Keep one compatible road-data version for the search, then check known closures before replying. Recompute a legal route or return the stated failure; a cached route needs the same check.

One road model; check closures before replying.

Return to lesson
Design maps and route planningWhich resources grow when route queries, map viewing or retained road versions increase?Recall first, then reveal

Route queries need CPU, map tiles need network bandwidth, and retained road graphs need memory.

Compute, bytes, versions.

Return to lesson
Design maps and route planningDoes the displayed ETA promise when the driver will arrive?Recall first, then reveal

No. It estimates arrival using the chosen road model and the updates received so far.

Model time is estimated time.

Return to lesson
Design event-time click analyticsC901 commits, but its queue acknowledgment is lost. Why does replay leave the count at 100?Recall first, then reveal

The processed-event record and count change committed together. Replay finds that record and skips the increment; if the transaction failed, neither change survived.

Save identity with the count; retries cannot count twice.

Return to lesson
Design event-time click analyticsA click arrives late. Which timestamp chooses its minute?Recall first, then reveal

When it occurred. Arrival and processing timestamps explain the delay.

Count when it happened.

Return to lesson
Design event-time click analyticsAn ad’s counts are split across workers. What must happen before ranking ads?Recall first, then reveal

Combine all partial counts for each ad, then compare the complete totals.

Sum before top.

Return to lesson
Design a distributed unique-ID generatorTwo threads request an ID from A’s range at once. What prevents a duplicate?Recall first, then reveal

The authority gives processes non-overlapping ranges. Inside A, an atomic cursor update or lock gives each thread a different value. Both checks are needed.

Separate ranges between processes; synchronize within each.

Return to lesson
Design a distributed unique-ID generatorA process crashes with unused IDs in its range. May its replacement reuse them?Recall first, then reveal

No. Leave that range unused and reserve a new one, accepting gaps.

Burn uncertain values.

Return to lesson
Design a distributed unique-ID generatorWhat does a unique ID leave for the application to decide?Recall first, then reveal

How to recognize a repeated business request, record creation time, check access and provide gapless numbering if required.

A key is not the whole contract.

Return to lesson
Design a webhook delivery platformNorthHarbor saves E402, but its 202 reply disappears. What changes on retry?Recall first, then reveal

Keep E402 and delivery D22; create a new transport attempt. NorthHarbor finds its saved inbox record, so repeated arrival need not repeat the local business effect.

Lost reply → same event, new attempt.

Return to lesson
Design a webhook delivery platformHow does the receiver stop a repeated event from repeating a local update?Recall first, then reveal

In one transaction, change business state and mark the verified inbox event DONE. A retry then sees that completion.

Effect and DONE commit together.

Return to lesson
Design a webhook delivery platformWhy can the sender not promise that every event finishes its business work?Recall first, then reveal

The receiver may stay unreachable, retention may expire, and completion depends on the receiver and any external provider’s guarantees.

Bound attempts; disclose outcomes.

Return to lesson
Design a feature-flag and configuration platformapp8 installs C18, then its slow C17 download finishes. Which version stays active?Recall first, then reveal

C18 stays active because installation accepts only a newer generation. A rollback also gets a new generation, even when it restores older values. Publication, installation and user exposure are separate events.

Newer generation wins, not the last download.

Return to lesson
Design a feature-flag and configuration platformHow can rollback restore older values without accepting an old download?Recall first, then reveal

Publish those values as a new, higher generation.

Old values, new generation.

Return to lesson
Design a feature-flag and configuration platformDoes downloading a cached snapshot prove it is current?Recall first, then reveal

No. Check the current generation with the publication service; cached bytes alone cannot show whether a newer version exists.

Ask which version is current; cached bytes cannot answer.

Return to lesson
Design a multitenant SaaS platformT7 and T8 both have invoice 17. What must every access path check?Recall first, then reveal

Verify the caller may act for the requested tenant. Keep that tenant in database relationships, cache keys, jobs, files and operator checks; the invoice number alone is insufficient.

Same invoice number; different tenant ownership.

Return to lesson
Design a multitenant SaaS platformHow does a maintenance move avoid two writable copies of one tenant?Recall first, then reveal

Stop source writes, finish admitted work, copy and verify the data, then enable only the destination. The move accepts a maintenance pause.

Stop writes → drain → copy and verify → enable destination.

Return to lesson
Design a multitenant SaaS platformWhen should a tenant receive dedicated infrastructure?Recall first, then reveal

When measured resource use disrupts others, or a specific storage location or administrative requirement justifies the cost.

Isolate for a concrete constraint.

Return to lesson
Design a distributed object storeU31’s completion succeeds but its reply is lost. What does retry return?Recall first, then reveal

Return the saved V5 result. Completion first verifies durable bytes, then commits the manifest, current-version pointer and result together. A retry does not publish another version.

Durable bytes → atomic publication → saved reply.

Return to lesson
Design a distributed object storeA new version is published during a download. Which bytes should the download use?Recall first, then reveal

Keep the same immutable version selected at the start, and read every part from it.

One read, one version

Return to lesson
Design a permission-aware RAG knowledge assistantThe taxi answer cites P7 v12 but drops manager approval. Which check catches the problem?Recall first, then reveal

Citation validation checks that P7 v12 was supplied; faithfulness checks whether the answer preserves its approval condition. Authorization separately checks whether the employee may use that passage.

Allowed evidence → valid reference → supported claim.

Return to lesson
Design a permission-aware RAG knowledge assistantWhat must be checked besides whether the answer includes a citation?Recall first, then reveal

Permission to use the passage, whether the citation identifies a supplied source, and whether that source supports the answer.

Allowed, identified, supported.

Return to lesson
Design a permission-aware RAG knowledge assistantSearch is unavailable. May the assistant say no supporting policy exists?Recall first, then reveal

No. It could not search reliably; that does not prove the evidence is absent.

Unavailable is not absent.

Return to lesson
Design an LLM inference platformTwo gateways see one free worker slot. Who decides which generation may start?Recall first, then reveal

The worker reserves its actual KV memory before starting work. Router reports can be stale. Admission also bounds input/output tokens and queue waiting, so request count alone is insufficient.

A capacity report is a hint; the worker reserves memory.

Return to lesson
Design an LLM inference platformThe client cancels while a GPU computation still uses its KV blocks. When can memory be freed?Recall first, then reveal

Stop new work, wait until running computations release the blocks, then free them.

Stop → drain → free.

Return to lesson
Design an LLM inference platformDoes a saved generation record let a new worker resume the exact computation?Recall first, then reveal

No. It does not contain live GPU KV memory or output events that were never saved.

Identity survives; execution may not.

Return to lesson
Design durable agent workflowsThe vendor creates O902 but its reply is lost. What may the replacement worker do?Recall first, then reveal

Recover the saved action using the same provider-supported key or a result lookup. If the provider cannot establish the outcome, keep it unknown for reconciliation rather than send a new purchase.

Lost purchase reply → recover the same action.

Return to lesson
Design durable agent workflowsWhy not use a fresh key after a purchase reply is lost?Recall first, then reveal

The first purchase may already exist. Recover its original action through the provider’s supported retry or lookup; a fresh key can create another order.

Unknown is a state

Return to lesson
Design a recommendation platformI11 is returned at position seven, but U7 closes the page before seeing it. Is that an impression?Recall first, then reveal

No. Record actual visibility separately from returned items. Retrieve a bounded candidate set, filter and rank it, then learn asynchronously from what the viewer actually experienced.

Returned is not seen; observation drives learning.

Return to lesson
Design a recommendation platformDoes including an item in the response count as an impression?Recall first, then reveal

No. Count the agreed event showing that the item was actually visible to the user.

Returned is not seen

Return to lesson
Design a live video-conferencing serviceThe signaling socket connects, but P2 hears nothing. Has the call succeeded?Recall first, then reveal

No. Signaling arranges membership and connection setup; media carries the conversation. Check usable received audio/video and the media network path, not just control connectivity.

Setup succeeds only when usable media arrives.

Return to lesson
Design a live video-conferencing serviceHow does an SFU reduce repeated browser uploads?Recall first, then reveal

Each publisher sends its selected quality streams once. The server forwards the appropriate versions to participants allowed to receive them.

Upload once, forward selectively

Return to lesson