System-design interview · Core interviews
Design a photo-sharing service
Separate durable photo publication, follower feed preparation and authorized image delivery, then scale each for its own workload.
You will learn to
- Trace original upload, verified variants and a follower feed read.
- Use image bytes and follower amplification to justify storage, CDN and hybrid feeds.
- Explain processing retries and why cached candidates cannot grant access.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Caching: cache hits, misses, write policies and invalidation · Database indexes: B-trees, composite keys and query access · Message queues, event logs, delivery guarantees, and backpressure · Data partitioning and sharding
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01An upload becomes a photo, then a feed candidate
Maya uploads a sunrise photo. The service should preserve the original, prepare a small preview, show the photo on Maya's profile and eventually include it in eligible followers' feeds. Those are related but separate results. Accepting bytes does not mean every image variant exists, and placing a photo ID in a feed does not grant permission to view it.
Support uploads, deletion, author galleries, following, private accounts, title search and a twenty-item home feed. Begin with chronological ordering; sophisticated recommendations and exact engagement counts are extensions. A private account requires approved followers. Define whether an already issued media link can remain usable briefly after access changes; a cached image cannot revoke bytes already downloaded.
A photo becomes published only when the variants required by its page exist. The uploader can see processing or failure while preparation continues. Feeds and search may lag publication, but they must not expose private or deleted content simply because an old candidate remains in an index. These decisions give us a small set of observable states before we choose queues or a CDN.
Before choosing storage or fanout, clarify whether the feed is chronological, how quickly an accepted photo must become ready and visible, and whether private-media access must be checked on every new request. The requirements below select that last policy rather than long-lived public image URLs.
02Functional requirements
Agree on these supported actions before selecting components.
Upload and publish photos. Authenticated authors upload an original, observe processing status and publish a photo only after its required preview and display variants exist.
Browse and discover. Provide author galleries, title search and a chronological home feed of up to twenty eligible followed-author photos.
Manage audience and deletion. Support follows and approval for private accounts. Owners delete photos; current access rules govern feed results and media delivery.
Recover interrupted work. Repeated upload completion and processing attempts preserve one photo identity and one accepted variant manifest.
03Non-functional requirements
Use these as illustrative interview assumptions to agree with the interviewer. Numerical targets require measurement; they are not claims about an existing product or a proven implementation. p95 (the 95th percentile) means 95% of measured operations finish within the stated time.
Workload. Use two million uploads/day and ten million feed opens/day: approximately 116 uploads/s and 579 feed requests/s at the fivefold peak. Preview delivery is approximately 10 TB/day; a fifty-million-follower author is a separate skew case.
Feed response time. Target regional p95 feed-metadata latency of 200 ms at the planning peak under normal operation. Image download and decoding are measured separately; this target is not a whole-screen rendering promise.
Readiness and freshness. For valid admitted images under the tested size and format limits, target p95 upload-complete-to-READY within 30 seconds. Target p95 propagation from READY to the prepared candidate list or author history used by an eligible active follower’s refresh within five seconds during normal processing. This makes the photo available for retrieval; the bounded twenty-item page need not display every new photo.
Durability. Accepted originals, published manifests and required variants must survive a worker restart or one storage-node failure within the region. A success status cannot depend on an encoder keeping its private local files.
Privacy. Check current audience permission before returning private photo metadata or admitting a new private-media transfer, including cache hits. Previously downloaded bytes and already admitted transfers are outside recall.
Isolation of expensive work. Bound upload bytes, decoded pixels, processing concurrency and follower batches. Large or malformed images and celebrity fanout must not exhaust ordinary feed-serving capacity.
04Trace one photo from upload to a follower
For a small service, use one application and a database holding photo metadata, follows and image bytes. The application receives the upload, validates its format and dimensions, produces a preview and a larger display image, then commits those bytes and metadata together. Return the photo identity after commit. If decoding fails, no published photo appears.
When follower Leo opens the feed, read the authors Leo follows, obtain a bounded recent set of their photos, check current visibility and sort by creation time with photo ID as a tie-breaker. Return twenty metadata records and the routes for their previews. The browser subsequently requests those image bytes. The author's gallery is simpler: query one owner's photos in order.
This baseline is intentionally small. It explains publication, feed assembly and media retrieval without making asynchronous jobs necessary for correctness. Its limits are easy to identify: image processing delays uploads, large bytes burden database backups and delivery, and each feed request repeats work across many authors. These limits motivate background processing, separate image storage and precomputed feed candidates.
The small baseline publishes bytes and metadata together, then serves a follower feed from eligible photos.
Read each connection in order
- syncUpload or open feedAuthor and viewer → Photo application
- syncPublish image set / query followed photosPhoto application → Photos, variants and follows
- syncStatus, metadata and image bytesPhoto application → Author and viewer
05Size image delivery separately from feed requests
Assume two million uploads/day, 200 KB average originals, one million daily active readers and ten feed opens per reader. Each page contains twenty 50 KB previews. That gives about 23 uploads/s and 116 feed requests/s on average; a fivefold peak is approximately 116 uploads/s and 579 feed requests/s.
| Resource | Estimate | Consequence |
|---|---|---|
| Original ingress | 2M × 200 KB = 400 GB/day | Bulk storage grows despite modest request rate |
| Ten-year originals | 400 GB × 365 × 10 = 1.46 PB | Retention, replicas and derivatives dominate bytes |
| Preview delivery | 10M pages × 20 × 50 KB = 10 TB/day | Media delivery needs a different capacity budget |
| Ordinary publication fanout | 23.1 photos/s × 300 active followers ≈ 6,930 inserts/s | Feed preparation creates extra writes |
A feed API returning IDs is not delivering all those preview bytes itself. A content delivery network, or CDN, can serve repeated immutable images near readers. At a measured 90% byte-hit rate, preview traffic reaching the origin would fall toward 1 TB/day, although viewers still receive 10 TB/day.
Average follower counts can conceal one celebrity. Fifty million follower references for a single photo are qualitatively different from three hundred. Estimate that separate case before choosing to write every publication into every feed.
06Separate upload acceptance from publication
Create an upload before transferring or publishing its bytes.
Start-upload request
POST /photo-uploads
| Request information | Purpose |
|---|---|
| Owner-scoped request key | Makes a retry recover the same upload. |
| Expected byte count | Lets completion verify the original's length. |
| Checksum | Lets completion verify the original's content. |
Response identities for the running example
Photo: p900
Upload: up900
A retry returns those same identities. Small deployments may carry bytes through the application; larger deployments can give the client a narrowly scoped direct-upload target.
| API | Meaning |
|---|---|
POST /photo-uploads/up900/complete |
Verify accepted bytes and begin processing |
GET /photos/p900/status |
Uploading, processing, ready or failed |
GET /feed?limit=20&cursor=... |
Eligible photo metadata and media routes |
GET /users/maya/photos |
Author gallery ordered by time and ID |
PUT /following/maya |
Follow or request approval for a private account |
DELETE /photos/p900 |
Owner revokes publication and starts cleanup |
Completion authenticates the owner and verifies the stored object; an upload URL alone cannot publish a photo. Return pending rather than claiming all previews already exist. A feed cursor carries its viewer and continuation context, so it cannot be reused for another person's private feed.
Title search can initially use a database text index. A later search service returns candidate photo IDs; the serving path still checks their current visibility. Search freshness and access permission are different contracts.
07Choose records around three access patterns
Photo metadata, upload state and feed references serve different purposes.
| Record | Stored information | Purpose |
|---|---|---|
Photo |
ID, owner, creation time, title, visibility, publication state, accepted image-variant references | Describes the photo and selects its published content. |
Upload |
Request identity, photo reference, accepted original | Connects a retryable upload to its photo and exact original. |
Follow |
Follower, author, approval status | Records the relationship and private-account approval. |
Feed |
viewerId, photoId, sortKey |
Stores a derived candidate reference, not another copy of the image. |
The photo page reads by ID. The gallery reads by (ownerId, createdAt, photoId). Publication fanout reads followers of an author, while feed assembly reads authors followed by a viewer. Those opposite relationship lookups need suitable indexes; a single primary-key lookup does not answer them all efficiently.
After moving bytes to object storage, a manifest records exactly which immutable original and variants belong to a published photo. It is a small metadata document, not the images themselves. Keep originals private and preserve their exact accepted identity. A path containing a photo ID does not automatically prevent later overwrites; use create-only writes or a pinned object version.
A durable outbox row records processing or publication work in the same transaction as the photo-state change. A worker later relays it. This closes the gap where metadata commits but a separate queue send fails, leaving a photo permanently stuck.
08Let slow image work continue without blocking the uploader
Move original bytes to object storage when their volume and backup cost justify the split. Reserve the upload, receive and verify the original, then commit PROCESSING together with a processing event. The owner sees that the original was accepted while the application remains free to serve other requests.
The image worker then performs these steps:
- Read the accepted original and decode it under pixel and memory limits.
- Create the preview and large variant.
- Write the outputs under immutable attempt-specific names and verify them.
- Ask metadata storage to publish their manifest.
Only a successful READY transition makes the photo eligible for feeds.
The queue may redeliver a job after a timeout. Repeating image work is acceptable; replacing a newer published result with a stale worker's output is not. The publication update must check that this worker still owns the active processing attempt and that the photo was not deleted. A repeated successful completion returns the accepted manifest. Unreferenced attempt outputs can be cleaned up once no valid attempt can publish them.
This adds jobs, storage operations and recovery work. Synchronous processing is still simpler when small bounded images reliably fit the response budget. The advantage of the queue is predictable request handling and resumable work, not automatic exactly-once execution.
09Trade repeated feed reads for recipient writes
The baseline gathers photos when the viewer reads. If Leo follows 500 authors and we inspect 100 photos per author, we consider 50,000 rows to return twenty. Repeating that work on every feed open can cost more than preparing a small list in advance.
For ordinary authors, a publication worker inserts the new photo ID into active followers' candidate lists. This is fanout on write: one publication creates many recipient references. Enforce uniqueness on (viewerId, photoId) so replaying a job does not create duplicate feed slots. Store progress between follower batches so a crash can resume without loading an enormous follower list at once.
For very popular authors, keep an author timeline and merge its recent photos during reads. This avoids millions of writes for followers who may never open the app. A hybrid feed combines prepared ordinary-author candidates with those popular-author timelines, deduplicates IDs, checks current eligibility and returns a bounded page. Choose the threshold from posting rate, active readership and measured read/write cost, not a universal follower number.
Current privacy checks remain necessary in both paths. Leo can unfollow Maya while a delayed worker still inserts p900. The stale reference may remain temporarily, but it must not authorize a private photo. Candidate preparation answers what to consider; access checks answer what may be returned.
Only the committed ready event enters fanout; retrying the insert keeps one candidate.
Read each connection in order
- syncPublish verified variant manifestImage worker → Photo metadata
- returnREADY committedPhoto metadata → Image worker
- asyncRelay ready eventPhoto metadata → Feed worker
- syncInsert viewer + photo if absentFeed worker → Viewer candidate list
- syncRepeat after uncertain responseFeed worker → Viewer candidate list
- returnSame candidate, no duplicateViewer candidate list → Feed worker
10Keep metadata decisions apart from media transfer
Place a CDN in front of immutable variants so a popular image can be served by many delivery locations. A feed request returns small metadata and approved media references; the image request retrieves bytes through the media path. Measure time to first preview independently of feed API latency.
For private photos, choose an authenticated media edge that checks the viewer's current access before serving bytes, including cache hits. This keeps authorization in front of delivery. A short-lived signed download grant is an alternative if access may remain valid until expiry; this design instead checks current permission. Do not call a URL viewer-bound unless the edge checks that viewer's session. Already downloaded bytes, and transfers admitted before a later revocation, cannot be recalled.
Replicate metadata according to the accepted-write durability requirement and protect originals through object-storage replication and backups. As metadata grows, distribute photo records while preserving author-gallery and viewer-feed indexes. Hashing photo IDs can spread stored rows but makes neither a gallery nor a feed local by itself.
Keep upload and feed-serving concurrency separately bounded. Slow large uploads should not occupy every worker needed for small feed reads. A ranking timeout may justify falling back to chronological order; a permission failure cannot justify serving unchecked private content.
For the chosen failure boundary, require durable majority commits across three metadata replicas in independent regional failure domains and object writes that survive one storage-node loss. Acknowledge original acceptance and publish manifests only after the relevant storage policy is satisfied. Measure queue capacity against the 30-second readiness and five-second feed-freshness targets.
The API accepts an original; image workers publish a verified manifest before feed workers distribute photo IDs. Feed reads merge those IDs with popular-author histories and check current access. Viewers fetch image bytes through the authenticated edge, which checks access even for cached variants.
Read each connection in order
- syncUpload / open feedAuthor or viewer → Upload and feed APIs
- syncState, histories and accessUpload and feed APIs → Photo metadata, follows + outbox
- mediaReceive originalUpload and feed APIs → Originals + image variants
- asyncRelay processing workPhoto metadata, follows + outbox → Image workers
- mediaRead original; write variantsImage workers → Originals + image variants
- syncPublish verified READY manifestImage workers → Photo metadata, follows + outbox
- asyncREADY events and follower pagesPhoto metadata, follows + outbox → Feed workers
- asyncInsert unique photo referencesFeed workers → Viewer candidate lists
- syncOrdinary-author candidatesUpload and feed APIs → Viewer candidate lists
- mediaRequest imageAuthor or viewer → Authenticated media edge
- syncCheck current private accessAuthenticated media edge → Upload and feed APIs
- mediaFetch variant on missAuthenticated media edge → Originals + image variants
11Test a crashed processor and delayed feed work
Worker W1 starts processing p900 and pauses after writing one preview. Another valid attempt completes the required set and publishes its manifest. W1 later resumes. Its publication must be rejected because its attempt is no longer current. Its independently named files cannot overwrite the accepted outputs. The owner sees one ready photo; unused files become cleanup work.
A feed worker inserts p900 for half a follower batch and crashes before saving progress. On retry it repeats that batch. Unique recipient/photo entries make the repeat harmless; checkpointing before completing the writes would instead miss recipients. Feed lag is visible in freshness metrics, while the authoritative published photo remains available on its page.
If the cache disappears, cap object-store requests and share one cache refill among concurrent requests for the same image so a popular image does not overload the object store. If metadata is unavailable, the media edge cannot establish current access and stops admitting new private transfers. Already delivered bytes cannot be recalled; a transfer admitted before a later revocation may finish.
Delete the metadata's published state first, then asynchronously remove feed references, search entries and obsolete media. Reads must still filter stale references. Physical deletion is background work, but the access decision cannot depend on every derived index having completed cleanup.
12Measure readiness, feed work and delivery cost
Monitor upload-to-ready delay, permanently failed decoding, processing backlog and missing manifest objects. For feeds, measure candidates examined per returned item, publication-to-feed lag, duplicate IDs and popular-author merge cost. For delivery, measure byte-hit ratio and origin traffic; a high request-hit ratio can hide expensive large-image misses.
Treat uploaded media as untrusted input. Bound encoded bytes, decoded pixel dimensions, CPU and memory. Restrict upload targets to the correct owner and object, and remove unnecessary location metadata according to product policy. Private media tokens and image contents should not appear in ordinary logs.
Storage cost includes originals, variants and retained failed-attempt outputs. Delivery cost includes CDN egress even when origin reads decrease. Feed preparation has its own write and memory cost: at 32 bytes per reference, 300 recipients use 9.6 KB per photo, whereas fifty million recipients use 1.6 GB before indexes and copies.
Test malformed images, response loss after upload acceptance, repeated processing, a deleted photo in a cached feed and a popular upload during cache loss. These tests examine the actual boundaries rather than whether each box on the architecture diagram is running.
13Check the design against the requirements
Use the agreed lists to check the finished design. The tests below still need to establish the targets; a proposed mechanism is not a measured result. FR refers to the numbered functional requirements above; NFR refers to the numbered non-functional requirements.
| Requirement | Design mechanism | Validation and remaining limit |
|---|---|---|
| FR1,4 + NFR3–4: ready photo | Immutable originals, durable processing work and current-attempt manifest publication. | Crash an encoder and resume an old attempt. Only complete verified outputs publish; benchmark the 30-second readiness target. |
| FR2 + NFR1–3: usable feed | Hybrid prepared candidates and popular-author histories, followed by bounded current checks. | Test ordinary and celebrity posts, measuring p95 feed response and READY-to-candidate propagation separately. Availability for retrieval does not guarantee a position in the twenty-item page. |
| FR3 + NFR5: private audience and deletion | Serving-time metadata and media-edge authorization; derived cleanup follows. | Remove access while old candidates and cached variants remain. New transfers must be denied; admitted bytes cannot be recalled. |
| NFR4: one-node survival | Replicated metadata plus accepted object writes under the chosen durability policy. | Fail a node after upload/READY acknowledgment and verify every referenced required object remains available. |
| NFR6: bounded cost | Separate feed, upload and processing budgets; bounded fanout batches. | Combine malformed images, slow uploads and a viral photo. Targets need measured spare capacity, not merely more workers. |
14Rapid revision
Remember: An uploaded original can survive while its preview fails. Publish only a verified, complete variant set.
| Path or choice | Purpose | Main limitation |
|---|---|---|
| Save the original before accepting | Preserve the source bytes | Accepted does not mean variants are ready |
| Processing and manifest | Publish the manifest of all required variants | Obsolete workers cannot replace the accepted manifest |
| Pull feed baseline | Assemble followed authors on demand | More followed authors increase per-read work |
| Prepare ordinary-author feeds; pull celebrity posts | Save repeated reads without millions of celebrity-post copies | Merge, deduplicate and check current access |
| CDN delivery | Serve repeated immutable image bytes near viewers | Check access and private-grant expiry |
| Partition records by photo ID | Spread independent records | Galleries and viewer feeds need their own indexes |
| Asynchronous cleanup | Remove obsolete derived data and files | Check permission even while obsolete entries remain |
In the interview, trace p900 from accepted original to READY manifest and then into Leo's feed. Use preview egress to justify delivery caches and follower amplification to justify the hybrid. End by showing how a stale worker and a stale feed reference are prevented from changing publication or access. Personalized recommendations can be added after those responsibilities are clear.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
When is an uploaded photo ready for a feed?
Reveal a model answer
After every required display variant exists and its accepted manifest is committed.
Interviewer follow-up
Why not publish when the original upload finishes?
Reveal the follow-up answer
Readers would encounter broken previews if decoding or resizing later fails.
What the answer must demonstrate: Separates original acceptance from published variants.
Why store processing work beside photo state?
Reveal a model answer
A durable outbox committed with PROCESSING avoids losing the job if the application crashes before a separate queue send.
Interviewer follow-up
Can the relay send twice?
Reveal the follow-up answer
Yes. Workers must identify the photo generation and make repeated completion harmless.
What the answer must demonstrate: Explains the database-to-queue crash gap and repeated relay.
Why not push every photo to every follower?
Reveal a model answer
Popular authors can create millions of writes for inactive readers. Pulling their recent timelines on demand can be cheaper.
Interviewer follow-up
How is the threshold chosen?
Reveal the follow-up answer
Compare posting rate, active followers, recipient-write cost and actual reader merge cost.
What the answer must demonstrate: Justifies the threshold with work, rather than a celebrity label alone.
Why check access after reading a feed cache?
Reveal a model answer
Candidate lists can outlive an unfollow, deletion or privacy change. Membership in that list is not permission.
Interviewer follow-up
Can cleanup replace the check?
Reveal the follow-up answer
No. Cleanup may be delayed and can race a late fanout insertion.
What the answer must demonstrate: Keeps authorization independent of derived-list cleanup.
How does a stale image worker fail safely?
Reveal a model answer
It writes attempt-specific outputs and must pass a current-attempt check before publishing the manifest.
Interviewer follow-up
Why are separate names necessary?
Reveal the follow-up answer
Metadata rejection would not help if the stale worker could overwrite bytes already referenced by the accepted manifest.
What the answer must demonstrate: Protects both metadata selection and immutable bytes.
Does a CDN eliminate image delivery cost?
Reveal a model answer
It reduces origin reads and can improve latency, but bytes delivered to viewers still consume bandwidth and cost money.
Interviewer follow-up
Which hit metric matters?
Reveal the follow-up answer
Byte-hit ratio, alongside request hits, shows how much origin traffic is avoided.
What the answer must demonstrate: Separates origin savings from total viewer egress.
Why is hashing photo IDs insufficient for galleries?
Reveal a model answer
A gallery selects one author’s photos in time order. Random ID ownership can scatter those records.
Interviewer follow-up
What supplies that access path?
Reveal the follow-up answer
An owner/time index or suitable author-based layout with an explicit hot-author policy.
What the answer must demonstrate: Matches indexes to query shape rather than primary storage alone.
What changes for a personalized feed?
Reveal a model answer
Add bounded candidate scoring and evaluate quality, while preserving publication and permission checks.
Interviewer follow-up
Can a high score override privacy?
Reveal the follow-up answer
No. Ranking orders eligible items; it cannot authorize them.
What the answer must demonstrate: Treats relevance and eligibility as separate decisions.
Blank-page exercise · 45 minutes
Build the answer yourself
Design photo uploads, galleries and a followed-author feed. Use 2M uploads/day and a celebrity with 50M followers to justify processing, storage and delivery choices.
- Use 5 minutes to agree numbered functional and non-functional requirements for publication, chronological feeds, readiness/freshness, durability and private access.
- Trace one upload and follower read in 8 minutes.
- Estimate original bytes, preview egress and fanout in 7 minutes.
- Define upload APIs, records and READY publication in 10 minutes.
- Explain hybrid feeds, retries and private media in 10 minutes.
- Use 5 minutes to check the final design against the numbered FR/NFR lists, name the latency/freshness load tests and explain remaining media-access limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a photo-sharing serviceMaya’s original is stored, but preview generation fails. Is the photo ready, and what should Maya see?Recall first, then reveal
It is not READY. Preserve the original and show processing or a failure; publish only after all required variants exist and their manifest commits.
Original saved does not mean photo ready.
Return to lessonDesign a photo-sharing serviceWhy prepare ordinary authors’ feed entries but fetch celebrity posts when a reader opens the feed?Recall first, then reveal
Preparation saves repeated reads for ordinary audiences. Pulling celebrity posts avoids writing a copy for millions of followers who may never read it.
Spend writes where readers benefit.
Return to lessonDesign a photo-sharing serviceA cached feed entry names a private photo. May the service return it immediately?Recall first, then reveal
No. It identifies a possible feed item; current permission checks decide whether the viewer may receive it.
Candidate first, permission next.
Return to lessonFinal revision
Summary and interview notes
Save the original, generate the required variants, then publish their complete manifest. Prepare feed candidates separately; check current access before returning photo metadata or bytes.
Remember these points
- Publish complete variants through a verified manifest.
- Record background work durably with metadata.
- Use hybrid feeds when recipient writes outweigh read savings.
- Check access before private metadata or bytes leave the service.
Interview tips
- Use the same photo ID through every stage.
- Calculate the celebrity case separately from ordinary followers.
Important qualifications
- Strict immediate media revocation and version-pinned direct upload protocols require additional detail.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Direct-upload version pinning
Needed to prevent a reusable upload capability from replacing an already accepted original.
- Processing lease and publication proof
Useful for reasoning through multiple delayed worker attempts.
- Identity-bound media authorization
Needed when signed media grants must be bound to an authenticated session.
- Partition and identifier alternatives
Useful when gallery locality, allocation or ownership migration becomes the limiting concern.
Technical references
- S3 presigned upload URLsExplains constrained direct-to-storage uploads; image publication and authorization remain application responsibilities.
- PostgreSQL multicolumn indexesDocuments why an index’s column order matters for the owner/time access path.
- Transactional outbox patternSupports a reliable publication-to-processing/feed handoff.
- Amazon S3: Retrieving Object VersionsExact version retrieval pins the accepted original even if a presigned URL later writes a newer version.
- Amazon S3: Conditional WritesAlternative create-only enforcement; an object-key naming convention alone does not prevent overwrite.
Practice marks stay in this browser.