System-design interview · Core interviews
Design a microblogging service
Build a chronological posting service, then use measured reader activity and follower skew to choose where feed assembly happens.
You will learn to
- Separate accepted posts from the feed views derived from them.
- Explain why ordinary authors and very popular authors need different feed strategies.
- Handle publication retries, partial fanout and changing visibility without duplicating or exposing posts.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Data partitioning and sharding · Caching: cache hits, misses, write policies and invalidation · Message queues, event logs, delivery guarantees, and backpressure
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Define publication and the home feed
Design a service where users publish short posts, follow authors, like posts and browse a chronological home feed. Replies reference a parent post; reshares reference an existing post while recording who shared it. Maya publishes “The bridge is open” as post p701. Leo follows Maya and should see p701 when he refreshes his feed.
For this exercise, limit text to 500 characters with an explicit Unicode counting rule. A successful publish means the post is durably stored and can be read from Maya’s profile. Leo’s feed may take a few seconds to include it. These are different completion conditions: storing a post does not mean every follower’s page has been updated.
Include deletion, private accounts and a bounded history of recent feed entries. Media uploads finish before a post references them. Begin with chronological ordering, not recommendations. Search, trends and notifications are separate consumers of published events; explain their boundaries if asked, but do not build them into the critical publication path. Global real-time ordering across all authors is also outside this interview’s contract.
Ask whether the home feed must be chronological or ranked, whether a post must appear to every follower before publish succeeds, and how private accounts behave. These choices determine the consistency and fanout work we actually need.
02Functional requirements
Agree on these supported actions before selecting components.
Publish and manage posts. Authenticated authors create and delete short posts, with verified media references and a 500-character limit under a declared Unicode counting rule.
Follow and browse. Users follow/unfollow authors, view author profiles and page through a chronological home feed of eligible followed posts.
React, reply and reshare. Support likes, replies linked to a parent and reshares linked to the original post. Deletion or privacy changes of the source still govern disclosure.
Recover repeated actions. Retries of one creation return its existing post; repeated like/unlike operations preserve the intended user/post relationship.
03Non-functional requirements
Use these as illustrative interview assumptions to agree with the interviewer. Numerical targets require measurement; they are not claims about an existing product or a proven implementation. p95 (the 95th percentile) means 95% of measured operations finish within the stated time.
Workload. Use 100 million posts/day and approximately 81,019 page requests/s at the fivefold peak. Include the fifty-million-follower celebrity case and media bytes rather than sizing from the average post rate alone.
Response time. Target regional p95 post-acceptance and feed-page latency of 300 ms under the admitted peak workload in normal operation. Media downloads are measured separately from the metadata page.
Feed freshness. Target p95 propagation from source commit to the prepared candidate list or author history used by an eligible active follower’s refresh within five seconds during normal processing. Publication does not wait for every follower; bounded retrieval may omit a post, and older pages may require refresh to consider late candidates.
Source durability. Accepted posts, retry results, relationships and publication work must survive a process restart or one database-node failure within the region. Prepared inboxes are rebuildable views, not the only source of a post.
Current access. Check current post visibility and relevant private-account relationships before returning content. Cached candidates and a ranking fallback cannot authorize a now-private or deleted post.
Bounded serving work. Cap candidates, follower batches and retained inbox windows. Choose incomplete but authorized pages or explicit overload responses over unbounded work; exact global ordering and exact engagement totals are outside the contract.
04Follow one post from creation to a reader
Start with one application and a replicated SQL database. Maya sends a create request with request key post-71. In one transaction the application inserts p701 and saves the request key’s result. It replies after commit. If the reply is lost, retrying post-71 returns p701 instead of creating a second post.
A profile request uses an index ordered by author and creation time. Leo’s home-feed request first reads the authors he follows, retrieves their recent posts and merges those ordered results. It returns the newest twenty eligible posts. This is fanout on read: one viewer request gathers candidates from several authors. Nothing needs to be copied into Leo’s account when Maya publishes.
Like actions use a unique pair of user ID and post ID. Repeating “like p701” leaves one relationship rather than incrementing an unprotected counter twice. A reshare stores p701’s identity instead of copying its body, so deleting the original can hide it everywhere that checks that reference.
This baseline is complete for modest traffic. Its weakness appears when many readers repeatedly merge the same author histories. We can measure that work before introducing background workers or precomputed feeds.
The application reads followed authors and their indexed posts, then returns an eligible chronological page.
05Estimate pages, candidates and media separately
Assume one billion registered users, 200 million daily users and 100 million posts per day. Each daily user opens two home-feed pages and five profile pages, with twenty posts per page. Use decimal bytes and a fivefold peak multiplier.
| Quantity | Calculation | Design implication |
|---|---|---|
| Post writes | 100 million / 86,400 ≈ 1,157/s | Publication is much smaller than reading work. |
| All page reads | 200 million × 7 / 86,400 ≈ 16,204/s | Peak is about 81,019 page requests/s. |
| Displayed items | 16,204 × 20 ≈ 324,074/s | Loading and checking records multiplies page traffic. |
| Text retention | 100 million × 310 B = 31 GB/day | About 56.6 TB over five years, before copies and indexes. |
| Media ingress | 20 million × 200 KB + 10 million × 2 MB | 24 TB/day: media dominates stored bytes. |
If Leo follows 200 authors and we fetch twenty candidates from each, one home page examines up to 4,000 candidates to return twenty. Across approximately 4,630 home-feed requests/s, that is roughly 18.5 million candidates/s. Profile reads have different costs and should not be included in that multiplication.
The opposite extreme is also expensive: copying a celebrity’s post to fifty million followers takes 1,000 seconds at 50,000 inserts/s. Average follower count hides this problem. Measure audience activity and the distribution of follower counts before choosing a fanout policy.
06Make retries and pagination explicit
The API exposes user intent while deriving the acting user from authentication. The request key belongs to one author and one payload; reusing it with changed text is an error.
Create-post request
POST /v1/posts
| Request information | Running example or meaning |
|---|---|
| Author-scoped request key | post-71 |
| Text | The bridge is open |
| Verified media IDs | References to uploads that have already completed, when media is attached. |
After commit, the running example returns post ID p701.
| Request | Meaning |
|---|---|
POST /v1/posts |
Create one post; return its stable ID after commit. |
GET /v1/users/u17/posts?before=cursor |
Read a bounded author history. |
GET /v1/feed?cursor=token&limit=20 |
Read eligible home-feed candidates in deterministic order. |
PUT or DELETE /v1/following/u17 |
Set or remove the follow relationship. |
PUT or DELETE /v1/posts/p701/like |
Set or remove one user’s like. |
DELETE /v1/posts/p701 |
Mark the owner’s post deleted. |
Use a cursor containing the last returned creation-time/post-ID pair and an initial upper cutoff. Numeric offsets shift when newer posts arrive. Keyset pagination avoids that shift, but it is not a frozen snapshot: a late fanout entry can still be missed until refresh. State that limitation. Current visibility checks run on every page even when its candidate list was prepared earlier.
07Keep source records separate from prepared feeds
A post is the source of its text, visibility and media references. An inbox entry is only a possible item for a viewer’s feed. Keeping that distinction allows the system to rebuild feeds without recovering post bodies from many inconsistent copies.
| Record | Important key or access path |
|---|---|
| Post | Primary post ID; author/time index for profiles and recent-author reads. |
| CreateRequest | Unique author/request-key pair with payload identity and resulting post ID. |
| Follow | Unique follower/author pair; indexes in both directions. |
| Like | Unique user/post pair; displayed count is a derived aggregate. |
| Inbox | Unique viewer/post pair, ordered by creation time and post ID. |
| Outbox | Durable publication or deletion event saved with the post transaction. |
An outbox is a database table of work that remains to be sent to other services. Saving its event in the post transaction closes the gap where p701 commits but the application crashes before notifying feed workers. A relay can send that event later and may send it more than once.
Choose author-owned database partitions for the scaled design. Post, request result, outbox and author/time index commit together within the owning partition. A post ID carries routing information or is resolved through a routing directory. Globally unique identifiers alone do not provide an efficient query for all posts by one author.
08Precompute the feeds that readers actually reuse
Add fanout on write for ordinary authors with active followers. A worker reads Maya’s followers in bounded pages and inserts p701’s ID into their inboxes. Leo’s next feed request becomes a short inbox range query followed by loading the matching posts. This exchanges background writes and storage for less repeated work during reads.
Do not apply that policy to every author. Very popular authors remain on the pull path: save their recent posts once and merge them into followers’ feeds at read time. The worked design is therefore hybrid. An ordinary post generates inbox references; a celebrity post does not initiate a massive follower scan. Determine the threshold from measured recipient-write costs and useful reader requests, not a famous fixed number.
The read service merges Leo’s inbox with recent lists of the popular authors he follows, removes duplicate IDs, loads post records in batches and checks current visibility. Bound candidates and source lists so an unusual follow graph cannot consume unbounded work. An incomplete but authorized page is preferable to a timeout caused by searching forever for twenty items.
Keep only a useful recent window in inboxes and rebuild a bounded window for returning inactive users. Store media in object storage and deliver it through an authenticated media edge backed by a content delivery network. The edge checks current access before serving cached private bytes. Already downloaded content cannot be recalled, and an in-flight transfer is not an instantaneous revocation guarantee.
Choose durable majority commits across three replicas in independent regional failure domains for each authoritative database group. This implements the one-node-loss target for source posts, relationships and saved publication work; candidate caches remain rebuildable. Benchmark both page latency and active-follower freshness before increasing admitted load.
Publishing commits the post and outbox event in an author-owned partition. Workers prepare ordinary-author inbox entries; reads merge them with popular-author histories and check current visibility. Media follows a separate authenticated delivery path so feed references alone cannot grant access to private bytes.
Read each connection in order
- syncPublish / read feedAuthors and readers → Post and feed APIs
- syncCommit / load / popular historiesPost and feed APIs → Author-owned post partitions
- syncCurrent follows and accessPost and feed APIs → Follows + access relationships
- asyncCommitted publication eventsAuthor-owned post partitions → Outbox relay + fanout workers
- syncPage ordinary-author followersOutbox relay + fanout workers → Follows + access relationships
- asyncUnique viewer / post insertsOutbox relay + fanout workers → Viewer inbox references
- syncPrepared ordinary candidatesPost and feed APIs → Viewer inbox references
- mediaRequest mediaAuthors and readers → Authenticated media edge
- syncCheck current private accessAuthenticated media edge → Post and feed APIs
- mediaFetch bytes on missAuthenticated media edge → Private media objects
09Recover partial fanout without skipping followers
Suppose the worker inserts p701 for viewer A, then crashes before viewer B. Retrying the entire page is safe because the inbox key is unique per viewer/post. A’s duplicate insertion changes nothing; B receives the missing entry. The worker records the next page checkpoint only after every insertion in the current page succeeds.
The checkpoint is durable progress, not a note that a worker fetched a job into memory. A repeated queue event resumes or repeats work using the same post identity. Use a conditional checkpoint update so two workers cannot overwrite newer progress with an older position. Workers may repeat the operation, but those repeats leave the same inbox entries. The queue need not deliver exactly once.
Follower relationships can change during the scan. Handle new follows with bounded backfill and check current follow/private-account rules while serving. The scan does not pretend to capture a globally frozen social graph. Likewise, changing an author between push and pull policies may temporarily produce the same candidate through both paths; ID deduplication makes that transition harmless.
If the fanout queue is behind, p701 remains visible on Maya’s profile while follower feeds lag. Add workers only while inbox storage has spare capacity. Unlimited consumers can turn a freshness problem into a database outage.
The checkpoint moves only after both viewers have their unique candidate entry.
Read each connection in order
- syncInsert viewer A / p701Fanout worker → Inbox store
- blockedCrash before viewer BFanout worker → Inbox store
- syncRetry A (same key), insert BFanout worker → Inbox store
- syncCommit next page checkpointFanout worker → Job progress
10Deletion changes eligibility, not just storage
Deleting p701 marks its source record deleted and records a cleanup event. Inbox, search and cache references may remain temporarily. They cannot be treated as permission to reveal the body. Before returning a candidate, load its current visibility and the membership needed for private access; if that check cannot be completed, omit or fail that item rather than guessing.
The same rule applies to reshares. A reshare can retain an actor and timestamp for internal history, but it cannot resurrect text from an inaccessible original. Removing stale references later reduces wasted work; it is not the mechanism that protects private content.
Like counts may briefly lag because they are computed from relationship changes. That is acceptable for a display count, while the unique user/post relationship answers whether Leo has liked p701. Do not reuse this approximate-count reasoning for authorization. Freshness, popularity and permission have different consequences when stale.
11Measure freshness and protect the source of truth
Track post acceptance latency, home-feed latency, candidates examined per page and oldest unprocessed fanout event. Measure the delay until an active follower can see a newly eligible post. A fast page showing yesterday’s content passes a latency check but fails a freshness goal.
Rate-limit publication, follows, likes and media upload separately. One small post can trigger millions of follower writes, so abuse limits must account for that work as well as request count. Validate media ownership and readiness before referencing it. Protect hot post records with batched reads and caches, but preserve the current-access check for disclosure.
During ranking or optional analytics outages, serve a chronological eligible feed. During permission-store failure, do not use that fallback to expose uncertain private posts. Database replicas, backups and restore exercises protect source posts and relationships; inboxes can be reconstructed from those sources.
Cost comes from media storage and delivery, fanout writes, repeated candidate merges and retained indexes. Compare those quantities for active audiences. A lower database query count is not automatically a cheaper system if it requires millions of never-read inbox entries.
12Check the design against the requirements
Use the agreed lists to check the finished design. The tests below still need to establish the targets; a proposed mechanism is not a measured result. FR refers to the numbered functional requirements above; NFR refers to the numbered non-functional requirements.
| Requirement | Design mechanism | Validation and remaining limit |
|---|---|---|
| FR1,4 + NFR4: durable publication | Author-local post, request result and outbox transaction; unique like relationships. | Lose publish replies and replay likes. Confirm one post/action and retained publication work after one node fails. |
| FR2 + NFR1–3: fast fresh feed | Ordinary-author fanout plus popular-author pull and a bounded merge. | Benchmark p95 response and five-second candidate propagation with celebrity skew and partial fanout failure. Bounded pages need not display every available post. |
| FR3 + NFR5: references obey source access | Reshare references and current post/relationship checks. | Delete or restrict the original while old inbox and reshare references remain. |
| NFR6: bounded continuation | Keyset cursor, candidate limits and bounded rebuild for returning readers. | Insert a late older candidate between pages; disclose that refresh may be needed, rather than claiming a frozen snapshot. |
13Rapid revision
Remember: A post commits once; follower copies can arrive later. Save fanout progress only after the inbox writes finish.
| Interview question | Chosen answer | Consequence to remember |
|---|---|---|
| What does publish success mean? | Post, retry result and delivery work committed together. | Follower feeds can still lag. |
| Why start with pull? | A working feed needs no background copies. | Repeated many-author merges become expensive. |
| Why use hybrid fanout? | Prepare active followers’ entries; fetch celebrity histories on demand. | Merge both sources and remove duplicates. |
| How do workers recover? | Unique viewer/post keys; save progress after writes. | Replaying a partial page is safe. |
| What controls deletion/privacy? | Current source visibility and membership checks. | Old feed entries do not grant permission. |
| How do pages continue? | Continue after the last creation time and ID. | Late candidates may wait until refresh. |
| What is the principal cost? | Media bytes plus push-versus-pull work. | Measure audience-size differences, not only the average. |
For a final explanation, walk p701 through commit, recoverable fanout, candidate merge and current visibility checking. Then name the tradeoff: the service accepts a few seconds of feed propagation delay to avoid a transaction across every follower. A stricter repeatable feed session or global publication order requires an additional contract and mechanism; neither is supplied merely by choosing a distributed database.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
How can the first version build a home feed?
Reveal a model answer
Read followed authors, query recent author-indexed posts, merge by time and ID, then filter visibility. This is complete without an inbox service.
Interviewer follow-up
What makes that expensive?
Reveal the follow-up answer
Many active readers repeat the same multi-author retrieval and comparison work.
What the answer must demonstrate: Explains the complete pull path and its repeated work.
Why not push every post to every follower?
Reveal a model answer
A celebrity can require millions of recipient writes for one post, including inactive viewers. Keep those author histories on a read-time merge path.
Interviewer follow-up
How would you choose the threshold?
Reveal the follow-up answer
Compare measured active audience reads, inbox-write cost, merge cost and freshness delay.
What the answer must demonstrate: Connects follower skew to hybrid fanout, not a universal threshold.
Why save an outbox event with a post?
Reveal a model answer
A crash between database commit and queue publication would otherwise lose the fanout trigger. The saved event can be relayed again.
Interviewer follow-up
Can the relay send duplicates?
Reveal the follow-up answer
Yes. Consumers use stable event and viewer/post identities rather than relying on one delivery.
What the answer must demonstrate: Identifies the commit-to-queue gap and repeatable consumers.
A worker crashes halfway through a follower page. What happens?
Reveal a model answer
Repeat the page; unique viewer/post keys suppress duplicate effects. Advance progress only after all writes complete.
Interviewer follow-up
Why not checkpoint when the page is fetched?
Reveal the follow-up answer
That would skip followers whose writes never occurred.
What the answer must demonstrate: Places checkpoint after effects and uses unique inbox keys.
Why does deleting every inbox entry not suffice for privacy?
Reveal a model answer
Cleanup may be delayed or incomplete. The read service must consult current post visibility and private membership before disclosing content.
Interviewer follow-up
Can an optional ranker outage bypass that check?
Reveal the follow-up answer
No. Recency is an ordering fallback, not permission to serve inaccessible content.
What the answer must demonstrate: Separates candidate freshness from authorization.
Does a keyset cursor create a frozen feed?
Reveal a model answer
No. It stabilizes the continuation position, but late fanout can add older candidates. The simple contract allows those items to appear on refresh.
Interviewer follow-up
What if repeated pages must use one fixed set?
Reveal the follow-up answer
Materialize a bounded candidate session and still recheck current visibility.
What the answer must demonstrate: States the keyset limitation before adding a session snapshot.
Why store likes as relationships?
Reveal a model answer
A unique user/post pair makes repeated like and unlike requests well-defined. Blind counter increments duplicate actions after retries.
Interviewer follow-up
Must the displayed count update synchronously?
Reveal the follow-up answer
Not for this scope; the displayed count may catch up as relationship changes are processed, while the saved user/post row tells us whether that user has liked the post.
What the answer must demonstrate: Distinguishes authoritative relationship from derived count.
Why distinguish page requests from item impressions?
Reveal a model answer
A page contains many records, media references and visibility checks. Multiplying by items per page exposes the actual serving work.
Interviewer follow-up
What further measurement changes the fanout choice?
Reveal the follow-up answer
Follower skew and active-reader reuse, because average follower count conceals celebrity amplification.
What the answer must demonstrate: Keeps page, item and fanout units separate.
Blank-page exercise · 45 minutes
Build the answer yourself
Design a chronological microblogging service. Trace one ordinary post and one celebrity post, including a lost publish response and a fanout worker crash.
- 0–5 min: agree numbered functional and non-functional requirements for chronological feeds, publication, five-second freshness, private access and source durability.
- 5–12 min: trace the SQL baseline and calculate page versus candidate work.
- 12–20 min: define post, follow, like, request and outbox records.
- 20–30 min: explain ordinary-author push and celebrity pull with a concrete cost comparison.
- 30–38 min: recover partial fanout and enforce deletion/privacy during reads.
- 38–45 min: review the final design against the numbered FR/NFR lists, test latency/freshness and failure boundaries, then state cursor, media-cost and ordering limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a microblogging serviceWhere is one published post stored, and how does it reach follower feeds?Recall first, then reveal
Save its source record and author-history entry, then create feed references through delivery work that can resume after a crash.
One body, many views.
Return to lessonDesign a microblogging serviceA worker inserts p701 for viewer A, then crashes before viewer B and before saving progress. What should the retry do?Recall first, then reveal
Repeat the follower page. The unique viewer/post key makes A’s insertion harmless, B receives the missing entry, and progress advances only after every page write succeeds.
Write, then checkpoint.
Return to lessonDesign a microblogging serviceA feed still contains a deleted or newly private post. What prevents disclosure?Recall first, then reveal
Before returning it, check the source post’s current visibility and the viewer’s current membership or relationship.
Candidates propose; permissions decide.
Return to lessonFinal revision
Summary and interview notes
Save one source post, then prepare feed references for active ordinary audiences and fetch celebrity posts on demand. Workers resume incomplete copying; readers still check current visibility.
Remember these points
- Start with an author-indexed pull feed.
- Use hybrid fanout to handle follower skew.
- Recover partial pages through unique inbox keys and write-before-checkpoint ordering.
- Check current access after candidate selection.
Interview tips
- Use the fifty-million-follower example to motivate the hybrid design.
- Keep page QPS separate from item and media work.
Important qualifications
- The basic cursor is not an immutable feed snapshot.
- Search, trends, strict regional recovery and global ordering remain separate extensions.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Overlapping workers and generation-aware checkpoints
Analyze stale workers and policy migrations beyond the basic retry example.
- Repeatable candidate sessions
Add a stronger browsing contract when late candidate arrival is unacceptable.
- Partition migration and regional recovery
Specify ownership transfer and recovery targets for larger deployments.
- Search, trends and recommendation algorithms
Build each derived product with its own ranking, freshness and abuse requirements.
Technical references
- PostgreSQL multicolumn indexesExplains why index column order matters for author/time and reader/time query shapes.
- Transactional outbox patternSupports durable post publication followed by recoverable fanout work.
Practice marks stay in this browser.