System designby Learnastra

Concept lesson · Foundations

Capacity estimation: throughput, latency, concurrency and storage

By Anup Rai

Start here

Definition

Capacity estimation translates an assumed workload into the compute, memory, storage and network resources needed to meet performance and failure targets. Throughput is work completed per unit time, latency is time per operation, and concurrency is work in progress.

Why it matters: Without a workload and units, “millions of users” cannot tell you how many servers or how much storage a design needs.

The visual modelCapacity estimates: request rate, bandwidth, and concurrency

Estimate request rate and bandwidth from the photo workload. Apply Little’s law separately at the metadata-service boundary.

Capacity estimates: request rate, bandwidth, and concurrencyEstimate request rate and bandwidth from the photo workload. Apply Little’s law separately at the metadata-service boundary. One million daily users each view twenty photos: twenty million views per day, about 231.5/s on average. The assumed ten-times peak is 2,315 views/s. At 100 KB per thumbnail it requires about 231 MB/s before overhead. Separately, a metadata service at 2,000/s and mean time 0.05 s has about 100 requests in flight in steady state. A daily user count is not a simultaneous connection count. State decimal bytes and the measurement boundary.Photo-service estimates: convert each unit explicitly1,000,000 users x 20 views/day = 20,000,000 views/daydivide by 86,400 secondsaverage = 231.5 views/sassume 10x peakpeak = 2,315 views/sPHOTO BANDWIDTH2,315/s x 100 KBabout 231 MB/sMETADATA CONCURRENCY2,000/s x 0.05 s100 in flight (mean)Little’s law uses means in a stable system. Photo bytes and metadata QPS are distinct loads.
Read the diagram step by step
  1. One million daily users each view twenty photos: twenty million views per day, about 231.5/s on average.
  2. The assumed ten-times peak is 2,315 views/s. At 100 KB per thumbnail it requires about 231 MB/s before overhead.
  3. Separately, a metadata service at 2,000/s and mean time 0.05 s has about 100 requests in flight in steady state.
  4. A daily user count is not a simultaneous connection count. State decimal bytes and the measurement boundary.

Worked example

One million users making twenty requests a day create 20,000,000 / 86,400 = about 231.5 requests/s on average. A stated 10x peak is about 2,315 requests/s; it is an assumption to validate, not something implied by the user count.

Key takeaways

  • Count requests, bytes and retained data separately.
  • Use peak load and surviving capacity when sizing.
  • Average concurrency = average arrival rate × average time spent in the same measured system, assuming stable operation.

You will learn to

  • Calculate average and peak rates with units.
  • Distinguish throughput, latency, concurrency, and storage.
  • Use an estimate to justify a design change.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: System design interview framework

Workload and timing examples are interview assumptions.

01Capacity estimation and its units

Capacity estimation translates a workload into the resources needed to meet its targets. A workload specifies what users do, how often, how large their requests are and how concentrated the traffic becomes. The estimate should be accurate enough to choose a design; a load test must later measure the actual implementation. Begin with four separate ideas. Throughput is completed work per unit time, such as 500 uploads per second. Latency is how long one operation takes. Concurrency is the number of operations in progress at once. Storage is how much retained data exists at a point in time.

A restaurant can serve many meals an hour while one customer's meal takes a long time. A batch service can likewise have high throughput and high latency. More concurrent work helps use idle resources; once the limiting resource is fully busy, extra work mostly waits. A claim of “10,000 users” needs to say what they do and when.

QPS means queries per second; in an API discussion people often use it for requests per second, so state whether you are counting API requests or database queries. Network capacity, often called bandwidth, is the maximum data rate a link or path can carry under stated conditions, usually measured in bits/s. Network throughput is the rate actually achieved. Requests/s × bytes/request estimates the required transfer rate; provision capacity above that demand, including protocol overhead and headroom. Peak means the busiest declared interval, while an average spreads all work over the entire measured period. These are different quantities even when one calculation produces another.

Quantity Example and unit Decision it informs What it cannot establish alone
Throughput 2,000 completed requests/s Required processing rate How long one user waits
Latency 50 ms per request on average Response-time objective Total sustainable traffic
Concurrency 100 requests in progress Connections and memory Whether queues are stable
Storage 73 TB retained originals Disk/object capacity Read/write operations per second
Required transfer rate 231 MB/s at peak Network and delivery path CPU cost of producing each byte

02Worked example: average and peak QPS

Define a workload before estimating resources. For this photo-service example, assume one million daily active users, 20 photo views and 0.1 uploads per user per day, a 2 MB average upload, a 100 KB thumbnail, and 1 KB of metadata per photo. Use decimal units: 1 KB = 1,000 bytes, 1 MB = 1,000,000 bytes, and one day = 86,400 seconds.

Quantity Calculation Approximate result
Uploads per day 1,000,000 × 0.1 100,000
Average uploads/s 100,000 / 86,400 1.16
Photo views per day 1,000,000 × 20 20,000,000
Average views/s 20,000,000 / 86,400 231.5
Assumed 10× peak 231.5 × 10 2,315 views/s

The calculation is a sequence: first count actions per day, then divide by seconds per day, then apply an explicitly assumed peak factor. For these inputs, 1,000,000 × 20 = 20,000,000 image views/day, 20,000,000 ÷ 86,400 ≈ 231.5 views/s average, and 231.5 × 10 ≈ 2,315 views/s peak. We size thumbnail delivery against the last rate, then verify it against measured bursts.

Worked example diagramThis chart follows the example’s rate calculation. A cold cache changes the last value from 116 to as much as 2,315 requests/s.
Capacity estimation: throughput, latency, concurrency and storage: architecture diagram1. 1M active users to 2. 20M thumbnail views/day: 20 views per user; 2. 20M thumbnail views/day to 3. 2,315 views/s assumed peak: divide by 86,400; then ×10; 3. 2,315 views/s assumed peak to 4. 95% hit cache: cacheable lookup; 4. 95% hit cache to 5. 116 origin misses/s: 5% miss fraction1 → 2: 20 views per user2 → 3: divide by 86,400; then ×103 → 4: cacheable lookup4 → 5: 5% miss fraction011M active users0220M thumbnailviews/day032,315 views/sassumed peak0495% hit cache05116 origin misses/s
  1. 1 → 220 views per user1M active users → 20M thumbnail views/day
  2. 2 → 3divide by 86,400; then ×1020M thumbnail views/day → 2,315 views/s assumed peak
  3. 3 → 4cacheable lookup2,315 views/s assumed peak → 95% hit cache
  4. 4 → 55% miss fraction95% hit cache → 116 origin misses/s

03Storage, retention and network bandwidth

The photo workload creates two different demands: storage for retained objects and network capacity for repeated delivery. Logical data counts one copy of each retained object; replicas and backups consume additional physical storage. Keep those counts separate from bytes sent to viewers:

Category Calculation Result and scope
New originals 100,000 × 2 MB 200 GB/day
One year of originals 200 GB/day × 365 73 TB before deletion, compression, indexes or redundancy
Three full copies 73 TB × 3 219 TB for originals alone
Metadata growth 100,000 × 1 KB 100 MB/day; 36.5 GB/year before indexes
Thumbnail delivery 20 million × 100 KB 2 TB/day
Average transfer demand 2 × 10^12 / 86,400 About 23.1 MB/s, or 185 megabits/s
Assumed 10× delivery peak 23.1 MB/s × 10 About 231 MB/s before headers and retransmissions

Keep these quantities separate:

  1. Count other stored data separately. Thumbnail variants and backups are additional categories; the replication multiplier does not include them.
  2. Bytes and metadata scale differently. Their large size difference is a reason to store them separately. The metadata database need not carry every byte transferred to viewers.
  3. Convert units explicitly. Network links are often rated in bits/s: multiply bytes by eight. State whether you mean MB or MiB.

04Little’s law and latency percentiles

Request rate alone does not tell us how many connections or request buffers are occupied. A request continues using some resources while it waits for storage or another service. To size those resources, relate the completion rate to the time each request remains in the service.

Concept in focusHow many requests are inside the service?

Each square represents one request. These are long-run averages for a stable service.

How many requests are inside the service?Each square represents one request. These are long-run averages for a stable service. Count five rows of twenty request squares inside the service. Arrivals and completions average 2,000 requests per second; mean time inside is 50 ms. Little’s law gives average in-flight work of 100, not a tail-latency prediction.A stable service: 2,000 requests/s; mean time 50 msarriveInside the service boundarycompleteEach square is one request: 100 in flight on average.L = 2,000/s x 0.050 s = 100. Use the same boundary and averages.

Remember: 2,000 requests/s x 0.050 seconds = 100 requests in flight.

Read the diagram
  1. Count five rows of twenty request squares inside the service.
  2. Arrivals and completions average 2,000 requests per second; mean time inside is 50 ms.
  3. Little’s law gives average in-flight work of 100, not a tail-latency prediction.
Try from memoryIf mean time doubles at the same stable throughput, what happens to average in-flight requests?

It doubles from 100 to 200: L = 2,000/s × 0.100 s. This assumes the service remains stable at that throughput.

If average time rises to 0.5 seconds while admitted traffic stays at 2,000/s, concurrency becomes about 1,000. The extra 900 requests need memory, sockets, and possibly database connections. An unbounded queue hides overload briefly while increasing latency. It does not create processing capacity.

The stable-system condition matters. If arrivals stay at 1,200/s while only 1,000/s complete, an unbounded backlog grows by 200 requests/s, or 12,000 requests in one minute. There is no steady finite average latency to insert into this calculation. Bound the queue and reduce admissions, or increase the bottleneck’s measured service capacity.

05Bottlenecks and failure headroom

A bottleneck is the resource that first limits the workload: for example, CPU, database writes or network transfer. Headroom is spare capacity reserved for bursts, uneven load and failures. Once a load test identifies the limiting resource, size enough instances to meet the target even with the chosen failures.

Assume a load test measures 800 requests/s per application instance while meeting the latency objective, and the target peak is 2,315 requests/s.

Concept in focusLosing one machine uses up the spare capacity

Each server block represents 800 requests/s at the measured latency target. The lower bar compares peak demand with surviving capacity.

Losing one machine uses up the spare capacityEach server block represents 800 requests/s at the measured latency target. The lower bar compares peak demand with surviving capacity. Remove one 800 requests/s block from the fleet and compare demand with what remains. Four instances supply 3,200 requests/s; three supply 2,400 requests/s. A peak of 2,315 uses 96.5% of surviving capacity, leaving 85 requests/s.Peak demand: 2,315 requests/sFour healthy instances800/s800/s800/s800/sOne instance lost800/s800/s800/s2,315 used / 2,400 surviving = 96.5%Only 85 requests/s of spare capacity remain at this tested limit.

Remember: Four servers can hide a problem that appears after one fails.

Read the diagram
  1. Remove one 800 requests/s block from the fleet and compare demand with what remains.
  2. Four instances supply 3,200 requests/s; three supply 2,400 requests/s.
  3. A peak of 2,315 uses 96.5% of surviving capacity, leaving 85 requests/s.
Try from memoryIs 2,400 requests/s enough for a peak of 2,315?

It covers the point estimate but leaves only 85 requests/s, about 3.5% of surviving capacity. That is little room for workload variance or measurement error.

Fleet Normal capacity Capacity after one loss Assessment
Three instances 2,400 requests/s 1,600 requests/s Almost no normal spare capacity; insufficient after failure
Four instances 3,200 requests/s 2,400 requests/s Little failure headroom for uneven load

The notation ceil(x) means the smallest whole number at least as large as x; a partial server cannot satisfy the remaining load. For a chosen maximum of 70% of tested capacity after one failure:

  1. Budget each survivor: 800 × 0.7 = 560 requests/s.
  2. Find the survivors needed: ceil(2,315 / 560) = 5.
  3. Add failure capacity: five survivors require six instances.

This is illustrative sizing, not a universal 70% rule. Real benchmarks, cost, autoscaling lag and failure domains determine the target.

Check downstream amplification

The database, network, and object store must support the same workload. Six application servers do not help if they all wait for one slow query. Estimate the read/write amplification: if each API call issues five database queries, 2,315 API calls/s can become 11,575 database queries/s.

Size CPU from CPU time

Compute demand has a different unit from elapsed latency. If a measured request uses 2 ms of CPU time:

  1. CPU demand: 2,315/s × 0.002 CPU-seconds = 4.63 CPU-seconds/s, about 4.63 fully busy cores.
  2. Utilization headroom: at a chosen 70% limit, ceil(4.63 / 0.7) = 7 usable cores before additional failure capacity.

06Cache working set and cost model

A cache stores copies of reused data. Its size depends on distinct hot entries, not total requests. Suppose 500,000 frequently viewed photo records occupy 1.4 KB each including key and bookkeeping overhead. That is about 700 MB per full cache copy. Ten million reads of those same entries do not require ten million stored entries.

A cache hit finds the requested value in the cache; a miss must fetch it from the underlying database or storage service, called the origin. The request hit rate is the fraction of cacheable requests served as hits. This rate turns the delivery estimate into an estimate of work still reaching the origin.

A 95% request hit rate reduces 2,315 cacheable lookups/s to about 2,315 × 0.05 = 116 misses/s under the same workload. But when the cache is empty, the origin can suddenly see all 2,315/s. Protect that origin and warm popular entries gradually. Track byte hit rate separately: a few missed large images may dominate bandwidth despite a high request hit rate.

For cost, write a symbolic model before using current provider prices: storage GB-month + read/write operations + delivered GB + compute time + replication/backup. A cheaper storage tier can have retrieval fees and slower access. The interview value is identifying the dominant cost and a way to measure it, not memorizing a vendor price that may change.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What is capacity estimation? Estimate QPS for one million users making ten requests a day.

Reveal a model answer

“Capacity estimation converts a workload into rates and resource needs. Here one million users × ten requests is ten million requests/day. Dividing by 86,400 seconds gives about 116 requests/s average. I still need peak concentration, bytes per request, latency targets and failure headroom before choosing server capacity.”

What the answer must demonstrate: Show denominator and units.

Foundation · Question 2

Can a system have high throughput and high latency?

Reveal a model answer

“Yes. A batch worker may finish thousands of items a second while each item waits minutes in a queue. Throughput describes the completion rate; latency measures one item’s elapsed time. I would measure queue wait and processing time separately.”

What the answer must demonstrate: Distinguish work rate from wait time.

Applied · Question 3

How much storage do 200 GB/day of uploads need after a year?

Reveal a model answer

“Without deletion, 200 × 365 is 73,000 GB, or 73 TB in decimal units. That is logical originals. I would separately add derived images, indexes, copies, and backups, then apply the retention policy.”

What the answer must demonstrate: Separate logical data from physical overhead.

Applied · Question 4

What happens at 2,000 requests/s if average latency grows from 50 to 500 ms?

Reveal a model answer

“Assuming both measurements cover the same system in stable operation, the average number of requests in progress grows from about 100 to 1,000. That can exhaust memory or connection pools even without a traffic increase. I would inspect downstream latency and bound admitted work.”

What the answer must demonstrate: Use seconds and matching averages.

Applied · Question 5

Should a cache hold 20% of yesterday’s requests?

Reveal a model answer

“Requests are not stored objects. I estimate distinct hot keys and bytes per entry. If a million requests hit one record, that is one cache entry. I use observed reuse and eviction behavior to choose the working set, then account for replication and overhead.”

What the answer must demonstrate: Count distinct retained entries.

Applied · Question 6

Three servers can just meet peak. Is that a resilient design?

Reveal a model answer

“Not if the requirement includes surviving a server failure at that peak. I calculate the remaining capacity after the failure and keep headroom for imbalance. If two survivors cannot meet the objective, I add capacity, reduce admitted work, or agree on degraded behavior.”

What the answer must demonstrate: Calculate surviving capacity.

Applied · Question 7

An API runs five database queries. Which QPS matters?

Reveal a model answer

“Count both. At 2,315 API requests/s and five queries per request, the database receives about 11,575 operations/s before retries or cache effects. I would check whether each query is needed, indexed and independent of the others.”

What the answer must demonstrate: Explain amplification rather than hiding it.

Applied · Question 8

A photo service serves 2,315 peak views/s at 100 KB each. Which measurements would change the storage or delivery design?

Reveal a model answer

“The stated peak is 2,315 × 100 KB = 231.5 MB/s, about 1.85 Gb/s before overhead. I would measure repeated-key reuse and permission constraints to evaluate a CDN, and measure metadata and CPU costs separately to decide where scaling helps. Peak QPS alone cannot determine daily delivered bytes or metadata growth; those require daily volume and stored bytes per upload.”

What the answer must demonstrate: Use a number to justify a decision.

Blank-page exercise · 15 minutes

Build the answer yourself

Estimate a file-sharing service with 2 million daily users, five 200 KB downloads each, and a 6× peak. Defend one architecture decision.

  • Compute average and peak requests/s.
  • Compute delivered bytes/day and peak bytes/s.
  • Explain one failure-headroom calculation.
  • Distinguish assumptions from measurements.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Capacity estimation: throughput, latency, concurrency and storageRate conversionRecall first, then reveal

Daily operations ÷ 86,400 gives average operations per second.

Actions → daily count → seconds.

Return to lesson
Capacity estimation: throughput, latency, concurrency and storageAt 2,000 requests/s and 0.05 seconds per request, how many are in progress?Recall first, then reveal

About 100 on average: 2,000 × 0.05. Little’s law requires a stable workload and matching measurement boundaries.

Rate × time = work in progress.

Return to lesson
Capacity estimation: throughput, latency, concurrency and storageCache sizeRecall first, then reveal

Distinct hot entries × bytes per entry, then copies and headroom.

Keys, not requests.

Return to lesson

Final revision

Summary and interview notes

Capacity estimates translate a declared workload into rates, retained bytes, concurrent work and resource demand. Size each component for its peak load and for the capacity it must retain after the failures you plan to tolerate, then validate the assumptions against a load test.

Remember these points

  • Average requests/s = daily requests / 86,400; a peak multiplier is a separate assumption.
  • Logical storage, replicas, derived objects, indexes and backups are separate physical categories.
  • Little’s law uses average arrival rate and average time for the same system in stable operation; substituting a latency percentile does not give average concurrency.
  • CPU-seconds per request differ from elapsed request time; both affect sizing in different ways.
  • A warm-cache miss rate is not the capacity requirement after cache loss.

Interview tips

  • Write units at every conversion, especially bits versus bytes and MB versus MiB.
  • Show the surviving capacity after the required failure, rather than counting only healthy servers.
  • End an estimate by naming the architectural decision it changes.

Important qualifications

  • The 10× peak and 70% utilization figures are example assumptions, not universal defaults.
  • If accepted requests keep arriving faster than they finish, the queue keeps growing. A stable-workload concurrency estimate no longer describes that overload.

Technical references

Practice marks stay in this browser.