System designby Learnastra

A path through the material

Learn, explain, then design.

Each stage builds ideas used by later lessons. Follow the path in order or use prerequisite links to repair a specific gap.

How the guide is organized

Concept chapters begin with a definition and visual model, then explain the mechanism, guarantees, trade-offs, and interview questions. Worked examples illustrate specific operations and failure cases. Each system design opens in the interview version: scope and clarifying questions, separate numbered functional and non-functional requirements, a working baseline, useful estimates, APIs and records, justified scaling, and critical failures. A closing table checks the design against its requirements before the separate rapid-revision table. Its Q&A and exercise use a 45-minute scope. The Advanced link at the top opens the preserved detailed lesson with stronger guarantees and deeper failure analysis.

These stages follow common interview frameworks, including Hello Interview’s delivery framework. Use estimates when they inform a decision and adapt the depth and sequence to the interviewer.

How to practise a lesson

  1. Learn the definition and mechanism. Then explain a worked example without looking.
  2. Answer the chapter’s questions aloud; include the reason behind each choice.
  3. Try the follow-up before revealing its answer.
  4. Do the timed exercise on a blank page. Write the functional and non-functional requirements first; finish by checking your design against them. Revisit the cards later.

1. Understand one request

Explain what the user can do, what counts as success, which data is needed, and how much traffic to expect.

  1. System design interview framework
  2. HTTP APIs and request lifecycle
  3. Capacity estimation: throughput, latency, concurrency and storage
  4. Distributed systems: scalability, reliability, availability and efficiency

2. Store and find the data

Trace a query and a write before distributing them.

  1. Databases, data models, and ACID transactions
  2. Database indexes: B-trees, composite keys and query access
  3. Storage engines and data models
  4. Load balancing: definition, algorithms and failover
  5. Caching: cache hits, misses, write policies and invalidation
  6. Proxies: forward proxy, reverse proxy and API gateway
  7. Data partitioning and sharding
  8. Consistent hashing and virtual nodes

3. Reason about concurrency and failure

Draw timelines for stale reads, concurrent changes, lost responses, and worker crashes.

  1. Replication and durability
  2. CAP theorem: consistency, availability, and partition tolerance
  3. Consistency models
  4. Transaction isolation
  5. Quorums, consensus, leases, and fencing
  6. Idempotency, retries, and timeouts
  7. Message queues, event logs, delivery guarantees, and backpressure
  8. Distributed transactions and sagas

4. Build production and modern foundations

Explain delivery, approximation, permissions, recovery, and observability.

  1. Real-time communication: polling, long polling, SSE, and WebSocket
  2. Probabilistic data structures
  3. Keyword search and vector retrieval
  4. Authentication, authorization, and tenant isolation
  5. Multi-region architecture and disaster recovery
  6. Production readiness: SLI, SLO, observability, and recovery

5. Practise complete systems

Rotate through different workloads. Then choose the remaining designs from the library.

  1. Design a URL shortener
  2. Design a chat messaging service
  3. Design a ticket-booking service
  4. Design a file synchronization service
  5. Design an API rate limiter
  6. Design a personalized news feed
  7. Design a payment system and ledger
  8. Design a permission-aware RAG knowledge assistant
  9. Design an LLM inference platform
  10. Design durable agent workflows

Practise adapting a design to a different prompt

A familiar diagram is a starting point. Identify the rule that changes before reusing its components. Choose one prompt for a 45-minute exercise; these variations do not add requirements to the original exercise.

VariationWhat changesWhere to practise
Hotel reservationsReserve room-type capacity on every occupied night. Enough capacity on one night does not guarantee the whole stay.Multi-night inventory example
Wallet transfersMove existing funds between two accounts while preventing concurrent transfers from spending the same balance.Internal transfer example
Service boundariesDistinguish adding application instances from giving components separate deployments and transactions.Monolith and service comparison

For any variation, restate the requirements, update the data model, trace one successful operation and one failure, then draw the resulting architecture. Explain which parts of the previous design still apply.

Specialist prompts need additional preparation

This library covers general product and infrastructure interviews. It does not yet contain full worked designs for the following specialist prompts. Related chapters provide useful building blocks, but do not replace their distinct requirements.

A 45-minute mock interview

5m
Clarify
7m
Estimate & model
10m
Trace a design
15m
Deep dive
8m
Failure & recap

This is a practice allocation, not a universal interview format. Adjust when the interviewer redirects you.

Assess your answerEvidence to look for
Agreed requirementsYou listed the user actions separately from measurable quality targets, stated assumptions and exclusions, and confirmed the important choices.
UnderstandableYou defined the terms and traced actual records.
GroundedYou used workload estimates and required queries to justify design choices.
Correct under failureYou showed the commit point, retry, and recovery result.
DefensibleYou explained a cost and a reasonable alternative.
CompleteYou checked the final design against the agreed requirements and identified targets that still need measurement.

Why the 2026 additions are here

The core skills remain practical reasoning, accuracy, reliability, and scalability, as described in Amazon’s interview guidance. The additional topics are a curriculum judgment informed by current production engineering, not a measured ranking of interview frequency.

Each lesson cites primary technical sources, reviewed in September 2026. Product features depend on their documented configuration; illustrative numbers are not current company measurements. Technology names are implementation options with stated trade-offs, not a requirement to select the newest release. Each chapter ends with summary points, interview tips, and qualifications. Dotted concept links open the relevant explanation in a new tab.

Read the complete book