Concept lesson · Foundations
Production readiness: SLI, SLO, observability, and recovery
Start here
Definition
Production readiness is the ability to operate a service reliably: measure user outcomes, detect failure, limit damage, deploy changes, and recover. An SLI (service-level indicator) is a quantitative measure of service behavior; an SLO (service-level objective) sets its target over a stated window. The error budget is the unreliability that target permits: for example, the allowed number of bad requests or the allowed downtime, using that SLO’s denominator and window.
Why it matters: A healthy process can still serve slow, incorrect, or incomplete results. Operators need measurements of the operations users depend on, such as uploading and viewing a photo, and tested procedures for recovering those operations after a failure.
An SLI measures service behavior, such as the fraction of photos ready on time. An SLO sets a target over a time window. Compare completion and error rates for the new version with the current version before expanding the rollout.
Read the diagram step by step
- For one million photo uploads evaluated in the rolling 30-day window and a 99.9 percent success SLO, at most 1,000 may miss the defined readiness deadline. Each of the ten equal blocks represents 100 allowed misses.
- P501 missing the deadline consumes one of the 1,000 permitted bad events, shown as one hundredth of the first block; 999 remain if it is the only miss.
- Metrics reveal the rate, logs identify a specific job, and traces locate time across stages.
- A canary compares the new version with the old before rollout expands.
Worked example
If 99.9% of one million accepted photos must become ready within 60 seconds, at most 1,000 may miss that target. A 200 response at upload time does not prove that background processing met the objective.
Key takeaways
- Measure whether the requested operation finishes correctly, including any required background processing.
- Metrics show the trend; logs and traces explain individual failures.
- A rollback, failover, or restore is complete only after the user-visible result is verified.
You will learn to
- Define a user-facing success indicator, objective, denominator, and time window.
- Use metrics, logs, and traces to distinguish a symptom from its cause.
- Explain a canary rollback and verify recovery without losing accepted work.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Distributed systems: scalability, reliability, availability and efficiency · Message queues, event logs, delivery guarantees, and backpressure
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01What is production readiness?
Production readiness means being prepared to run the service through ordinary traffic, changes, overload, and failures. First define what users must be able to do and how reliably and quickly the service must respond. Then decide how to measure whether it meets those requirements and how to recover when it fails. Observability is the ability to understand internal behavior from the metrics, logs, and traces that the service produces.
An SLI (service-level indicator) is a quantitative measure of service behavior, such as the fraction of photos ready within 60 seconds. An SLO (service-level objective) is a target for that measurement over a window. An SLA (service-level agreement) is a commitment with agreed consequences, often contractual; it is not simply another name for an internal SLO. An error budget is the amount of failure the SLO permits over its measurement window, such as the number of requests allowed to miss a completion deadline.
Measure whether the service finishes the operation the user requested, including any background processing required before the result is usable. In the example, upload P501 is accepted immediately, spends 95 seconds queued, takes four seconds to render and one second to publish, and becomes ready after 100 seconds. That event misses a 60-second completion threshold despite a successful acceptance response and running API processes.
Check the complete user operation, protect the resources it needs, deploy changes safely and test recovery. A running process is useful evidence, but does not prove the service is fast enough, saves data correctly or enforces access permissions.
02SLI, SLO, SLA, and error budget: definitions and calculation
A service-level indicator, or SLI, is the measured behavior. A service-level objective, or SLO, is its target over a stated window. For completion, the numerator counts eligible photos ready within 60 seconds. The denominator counts eligible accepted photos whose evaluation period has elapsed. A just-accepted photo cannot be labeled late before its 60-second allowance ends.
This bar represents the 1,000 permitted bad events. It does not represent all traffic.
Remember: Allowed bad events minus actual bad events gives remaining budget.
Read the diagram
- Split a 1,000-event error budget into used and remaining portions.
- A 99.9% target over 1,000,000 eligible events permits 1,000 bad events.
- 400 bad events use 40% of that budget, leaving 600; the measured good fraction is 99.96%.
Try from memoryHow many additional bad events fit in the current fixed window?
600, assuming the window still contains exactly 1,000,000 eligible events and the target remains 99.9%.
| Decision | Example metric |
|---|---|
| Operation | Valid uploaded photo becomes viewable |
| Good event | Ready no later than 60 seconds after acceptance |
| Denominator | Eligible accepted photos with an elapsed evaluation period |
| Target and window | At least 99.9% over rolling 30 days |
| Separate guardrail | Valid upload attempts are accepted successfully |
Calculate the error budget
For one million evaluated photos, the 0.1% allowance permits at most 1,000 bad completion events. This allowance is an error budget. P501 is one bad completion event that consumes this budget; one late photo alone does not prove the aggregate 30-day 99.9% SLO was violated. Specify whether unsupported file types, canceled uploads, and failures caused by our service count. Exclusions should reflect the contract, not hide inconvenient incidents. Availability, timely completion, and correctness can require different indicators.
Define exactly which events enter the window
Handle no traffic and missing telemetry
When there are zero eligible events, the ratio is undefined, not 100% healthy. Use a no-data signal and the separate acceptance indicator or synthetic check. A time-based 99.9% availability target over 30 days permits 43.2 minutes of bad time, but that is a different denominator from the one-million-photo event budget. Do not convert between them without traffic assumptions.
A synthetic check performs a controlled test operation, such as uploading a test image and verifying that it becomes viewable. It can reveal a broken path when real users are inactive. Report that test separately from the real-user completion ratio rather than using it to invent a denominator for a no-traffic period.
03Observability and the four golden signals
Metrics are numerical measurements over time. For this service, track upload demand, timely completion, queue age, worker capacity, and errors. The classic four signals are latency, traffic, errors, and saturation: how long work takes, how much arrives, what fails, and which resource is nearly full. Google SRE monitoring.
The 100-second completion is the symptom. High queue age tells us where to investigate; it is not yet the cause. CPU may be low because a worker-concurrency setting is too restrictive, not because there is no demand.
| Observation | What it tells us | What it does not prove |
|---|---|---|
| Upload responses succeed | Acceptance path is responding | Photos become ready promptly |
| Queue age rises | Work is waiting longer | The queue service is broken |
| Worker CPU is 25% | CPU is not fully occupied | Sufficient workers are active |
| New-release cohort is slower | Release is a useful suspect | Causation without further inspection |
Break down metrics by processing stage and software version. Control label cardinality: the number of distinct label values and combinations that create separate time series. A separate time series for every photo ID would be costly; IDs belong in targeted event records and traces.
04Logs and distributed traces: locate the missing 95 seconds
Logs record individual events; structured fields make those records searchable. Traces connect work across stages so we can follow one request or asynchronous job. A span records one timed operation within a trace, such as a database call or a worker processing a photo. Carry a correlation identifier that links records for the same job without exposing secrets or personal data through acceptance, queue delivery, rendering, and publication. For asynchronous work, preserve the relationship even when it is represented by a trace link rather than one continuously open call.
Bar width is elapsed time. Child spans overlap the parent’s time and must not be added to it.
Remember: Read the timeline to locate the slow segment.
Read the diagram
- Locate the database and remote-call durations inside a 100 ms API span.
- The database call runs from 10 to 30 ms. The remote call runs from 35 to 90 ms.
- Other work or waiting occupies the unlabelled intervals.
Try from memoryShould the API time be calculated as 100 + 20 + 55 ms?
No. The child spans occur inside the 100 ms parent interval; adding them double-counts their time.
| P501 event | Elapsed time | Evidence |
|---|---|---|
| Upload accepted durably | 0 seconds | Acceptance record |
| Worker begins | 95 seconds | Queue/job trace |
| Rendering finishes | 99 seconds | Worker span or event |
| Photo becomes ready | 100 seconds | Publication record |
Rendering took four seconds and publication one. Almost all delay was before work began. We inspect the new worker release and discover that its concurrency limit was unintentionally reduced. That mechanism fits both the queue wait and low CPU.
Logs must not copy private image contents, access tokens, or unnecessary personal data. A photo ID and authorized diagnostic lookup are usually more useful than dumping the entire payload into an unrestricted log.
A practical implementation can instrument request and worker spans with OpenTelemetry, propagate trace context in the job metadata, and export selected traces and structured logs to a backend. Use the durably stored job record to decide whether a job completed; sampled traces are diagnostic evidence, not a complete SLO denominator. Cross-host timestamps may differ, so record stage durations with monotonic timers, which measure elapsed time without jumping when the system clock is adjusted, and account for clock uncertainty when subtracting timestamps from different machines.
- 1 → 2valid authenticated uploadUpload P501 → API: accepted at 0s
- 2 → 3stored processing jobAPI: accepted at 0s → Queue: wait 95s
- 3 → 4worker begins at 95sQueue: wait 95s → Worker: render 4s
- 4 → 5render done at 99sWorker: render 4s → Publish: 1s
- 5 → 6ready at 100sPublish: 1s → Outcome: ready at 100s
- 6 → 7completion misses 60s objectiveOutcome: ready at 100s → Trace + release evidence → action
05Actionable alerts and error-budget burn rate
A dashboard helps investigation; an alert asks someone to act. Paging on every brief CPU spike creates noise and does not necessarily protect the completion objective. Tie urgent alerts to significant user-impact or rapid budget consumption, with enough evidence to identify the affected service and likely response.
Suppose a recent window has 2% late photos while the SLO allows 0.1%. The burn rate is 2% / 0.1% = 20: the service is consuming its error allowance at twenty times the reference rate under that measurement. Use both shorter and longer windows so a severe ongoing problem is detected without treating a tiny transient sample as a sustained incident. SLO alerting reference.
Also monitor correctness constraints. A timely response that exposes a private photo is not a successful product outcome. Audit access-control decisions and check that rules such as “only authorized users can view a private photo” hold; latency metrics cannot establish confidentiality. The security-and-multi-tenancy chapter explains where and how to enforce those access checks.
For the illustrative 30-day window, a sustained 20× burn would consume a full window's budget in about 30 / 20 = 1.5 days under steady traffic and the same bad-event definition. That is a planning approximation, not a promise about a rolling window with changing request rates. Each paging alert should identify the affected objective, the team responsible for responding, a link to diagnostic information, and the first safe action to reduce the impact. Route slower budget erosion to a nonurgent work queue rather than paging on every symptom.
06Canary deployments, rollback, and backlog recovery
A canary release sends a limited portion of work to a new version before broad rollout. Compare workers running the new version with a control group running the current version on similar jobs. Measure whether photos become ready on time as well as whether the worker processes are running. In this example, route comparable jobs to a small canary worker pool with its own bounded queue so queue wait can be attributed to that pool. The canary shows elevated waiting and the reduced concurrency setting; stop expansion and restore the known-good configuration. If old and new workers instead pull from one shared queue, queue age is a shared symptom, not a per-version causal measurement. Compare per-version processing throughput and controlled workload evidence before attributing the delay.
Recover work accepted during the rollout
Keep data formats compatible
A schema change may prevent a simple binary rollback if the old code cannot read new data. Deploy changes in stages so that old and new application versions can both read the stored data during the transition. Feature flags can enable a new behavior separately from deploying the code. Limit the blast radius—the number of users or resources affected by one mistake—through gradual deployment and workload isolation. Keep a clear incident record of the symptom, change, action, and measured recovery.
Blue-green deployment: switch environments
A blue-green deployment prepares a second application environment, validates it, then shifts traffic from the old environment to the new one. It gives a clear traffic rollback target, but temporarily duplicates capacity and still needs connection draining: stop sending new work to the old environment while allowing its existing requests or connections to finish. A canary instead exposes a bounded cohort to the new version before broader rollout; either pattern needs comparable outcome measurements.
What traffic rollback cannot undo
07Disaster recovery: RPO, RTO, failover, and restore
A lost worker can be replaced and its jobs redelivered. A lost region may require a wider failover. A replicated bad deletion may require restoring older history. Choose the response from the actual failure, rather than treating every incident as a request to restart machines.
Recovery point objective, RPO, describes the acceptable loss of recent data measured in time. Recovery time objective, RTO, describes the target time to restore useful service. Both require tested procedures and measured results. For accepted photos, verify that the original files, records of pending processing jobs, publication status, and access permissions all survive recovery. Restoring one database does not by itself prove that users can upload and view photos again. Recovery guidance.
The multi-region-and-disaster-recovery chapter develops region placement and failback. Here the operational lesson is evidence: rehearse the recovery, check the customer-visible result, and record whether the objectives were met. A successful backup command or green failover control-plane status is only partial evidence.
Recovery planning also identifies the team responsible for recovery, a runbook with step-by-step instructions, accessible credentials and keys, and the dependencies needed to serve the recovered data. Verify that the remaining system can handle the required load when a server, zone, or region covered by the recovery plan is unavailable, and perform a restore to an isolated environment before relying on the procedure. Recovery point is a target: asynchronous replication lag must be measured to determine whether the observed lost work meets that target. Backups that share the same destructive permissions and retention policy as live data can fail together.
08Interview answer: explain how you know a service is healthy
Interviewer: “How will you know the upload service is healthy?”
Candidate: “I would measure both valid upload acceptance and whether accepted photos become ready within the agreed time. P501 returned success immediately but took 100 seconds, so an HTTP-success dashboard would miss the completion failure.
“I would trace acceptance, queue wait, rendering, and publication. The 95-second wait points toward processing capacity, and the canary’s reduced concurrency setting explains it. I would roll back that setting, verify the backlog drains, and check that replayed jobs preserve one correct result and private access. Alerts would focus on completion failures and error-budget burn.”
This answer connects monitoring to action: what the user needed, which measurements distinguish likely causes, what change is safe to undo, and how to check recovery. A monitoring box in a diagram needs those explanations.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Define SLI, SLO, SLA, and error budget with one user-visible example.
Reveal a model answer
An SLI (service-level indicator) is a measured user-visible behavior. An SLO is its target over a window. An SLA is an agreement about service commitments and consequences, often contractual. An error budget is the amount of failure permitted by the SLO.
For a photo service, measure the fraction of eligible accepted photos that become ready within 60 seconds, after each photo's evaluation period has elapsed. Set an illustrative SLO of at least 99.9% over 30 days. Of one million evaluated photos, at most 1,000 can miss the target. A separate SLA could specify a contractual commitment and credits; it need not use the same threshold as the internal SLO. Also measure acceptance so the service cannot make its completion ratio look good by rejecting every upload.
Interviewer follow-up
Why wait for the evaluation period to elapse?
Reveal the follow-up answer
“A photo accepted five seconds ago has not yet missed a sixty-second deadline. I evaluate it once when that deadline passes and count the durable ready-by-deadline result. The rolling window uses those evaluation times; unfinished jobs and missing telemetry must not disappear from the denominator.”
What the answer must demonstrate: A percentage without a denominator and window is incomplete.
Could accepting no uploads make your completion SLO look perfect?
Reveal a model answer
“Yes, if it treats no data as success, or shows only accepted uploads and hides rejected attempts. Zero evaluated photos proves nothing about completion. Also measure how many valid attempts are accepted, signal missing data and use a test upload when useful. Then we can distinguish failure to accept uploads from failure to process them.”
Interviewer follow-up
Should malformed uploads count as service failures?
Reveal the follow-up answer
“That depends on the specified contract, but I would separate expected validation rejection from failures of valid requests and avoid exclusions that hide our defects.”
What the answer must demonstrate: Beware metrics that improve by refusing useful work.
For 1,000,000 evaluated operations and a 99.9% success SLO, what is the error allowance? What burn rate does a 2% bad-event rate represent?
Reveal a model answer
“At 99.9%, one million evaluated operations allow one thousand bad events. If a recent window has two percent bad events against a 0.1 percent allowance, its burn rate is twenty. I would interpret that with traffic and window size before deciding how urgently to page.”
Interviewer follow-up
Why use more than one alert window?
Reveal the follow-up answer
“A short window detects rapid deterioration; a longer one helps establish that it persists. The combination reduces both slow detection and noisy reaction to tiny samples.”
What the answer must demonstrate: Keep percentage points and ratios distinct.
A photo takes 100 seconds: 95 queued, 4 rendering, 1 publishing. What does this trace reveal that low CPU usage does not?
Reveal a model answer
“The trace assigns ninety-five seconds to waiting, four to rendering, and one to publication. Low CPU cannot tell me whether concurrency is accidentally restricted or demand is absent. The trace locates the delay; release/configuration evidence then helps identify the cause.”
Interviewer follow-up
Does a high queue age prove the broker is faulty?
Reveal the follow-up answer
“No. Slow or insufficient workers can produce the same symptom. I would inspect service rates and stage behavior rather than blame the queue by its name.”
What the answer must demonstrate: Separate symptom, location, and causal evidence.
When do you use metrics, logs, and traces?
Reveal a model answer
“Metrics show aggregate trends and support alerts. Structured logs record individual events. Traces or correlated job events connect a specific journey across stages. For P501 I use metrics to detect late completion and the trace plus targeted logs to explain where it waited and which release handled it.”
Interviewer follow-up
Should photo IDs be labels on every metric?
Reveal the follow-up answer
“Usually not. That creates unbounded time-series cardinality. Keep per-photo details in appropriately protected logs or traces.”
What the answer must demonstrate: Choose the evidence type according to the question.
What should the canary compare before full deployment?
Reveal a model answer
“Compare similar workloads, user-visible completion, throughput and stage delay, with enough observations to distinguish a signal from noise. For worker changes, separate canary and control pools can make queue-wait attribution meaningful. If both versions share a queue, rising age affects the cohort comparison and cannot by itself blame one version. I would inspect per-version throughput/configuration and stop expansion or revert when the canary violates the agreed guardrails.”
Interviewer follow-up
Can every deployment be rolled back by restoring old binaries?
Reveal the follow-up answer
“No. Incompatible data/schema changes may make old code unsafe. I would plan compatible transitions and a recovery path before rollout.”
What the answer must demonstrate: Deployment safety includes data compatibility.
The old worker version is back. Can you close the incident?
Reveal a model answer
“Only after verifying the backlog drains and timely completion recovers. I also check that retrying a job did not publish duplicate results or repeat other side effects, and that photo permissions remain correct. Restoring the old version is a recovery step; I still need to verify that users can upload and view photos successfully.”
Interviewer follow-up
What if arrival rate still equals processing capacity?
Reveal the follow-up answer
“Existing backlog will not drain. I need temporary spare capacity or reduced admission and must communicate the ongoing delay.”
What the answer must demonstrate: Verify recovery under continuing load.
How do RPO and RTO change your recovery exercise?
Reveal a model answer
“RPO tells me how much recent accepted work may be lost; RTO tells me how soon useful service should return. I would measure both during a drill and verify original files, records of pending jobs, publication status, and permissions, rather than timing only a database restore command.”
Interviewer follow-up
Does replication replace that exercise?
Reveal the follow-up answer
“No. Replicas can carry the same mistake, and recovery has dependencies beyond copying state. We need to verify that uploads, background processing, and authorized viewing all work afterward.”
What the answer must demonstrate: Recovery objectives apply to the service outcome.
Blank-page exercise · 18 minutes
Build the answer yourself
Design a dashboard and incident response for P501 becoming ready at 100 seconds despite a successful upload response. Compare a canary worker release with the control.
- Define eligible requests and separate acceptance from timely completion.
- Calculate the error allowance and burn-rate example.
- Use a trace to identify where the 100 seconds was spent.
- Describe rollback, backlog recovery, and a check that private photos remain private.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Production readiness: SLI, SLO, observability, and recoveryWhat does an SLI measure?Recall first, then reveal
An actual service behavior, such as the fraction of accepted photos ready within 60 seconds. The SLO is the target and evaluation window.
Indicator measures; objective targets.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryHow many bad events does 99.9% permit among one million evaluated photos?Recall first, then reveal
At most 1,000 for that defined metric and window.
One in a thousand is the allowance.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryWhy is P501 a bad completion event despite HTTP success?Recall first, then reveal
Acceptance completed, but the photo waited 95 seconds and became ready at 100 seconds, beyond the 60-second good-event threshold. It consumes error budget; the aggregate SLO depends on all evaluated events.
Upload accepted ≠ photo ready.
Return to lessonProduction readiness: SLI, SLO, observability, and recoveryWhen is rollback actually successful?Recall first, then reveal
The known-good configuration is restored, queued work is draining, photos become ready on time, and their contents and access permissions are correct.
Changed back is not yet recovered.
Return to lessonFinal revision
Summary and interview notes
Production readiness means setting measurable reliability targets, limiting how many users a faulty release can affect, and testing recovery procedures. A running process or successful rollback command does not prove recovery: users must again be able to complete their operations, queued work must drain, and data and permissions must remain correct.
Remember these points
- An SLI (service-level indicator) is a measurement, an SLO is its target and window, and an SLA is an agreement with consequences.
- A 99.9% event SLO over one million evaluated photos allows 1,000 missed outcomes; zero events supplies no success evidence.
- Check each upload once when its readiness deadline arrives, including uploads still unfinished. Counting only completed jobs hides stuck work.
- A 2% bad-event rate against a 0.1% allowance is 20× burn, interpreted with traffic and window size.
- Metrics identify impact; traces and logs investigate causes; controlled canary evidence supports a release decision.
Interview tips
- Write the denominator, deadline, exclusions, and rolling-window rule before drawing a dashboard.
- Separate acceptance, timely completion, correctness, and confidentiality instead of treating an HTTP success response as proof of all four.
- For a worker canary, ask whether shared queues and workloads make the cohorts comparable.
Important qualifications
- Sampled traces cannot stand in for a complete SLO event counter; missing telemetry needs detection.
- A duration budget and an event budget are different measures even when both use 99.9%.
- RPO and RTO are objectives to demonstrate in a drill, not guarantees created by configuring replication or a backup job.
Technical references
- Google SRE: Monitoring Distributed SystemsPrimary operational reference for latency, traffic, errors, saturation, and useful monitoring.
- Google SRE Workbook: Alerting on SLOsPrimary reference for error-budget burn and multi-window alerting.
- AWS Well-Architected: Plan for Disaster RecoveryVerified reference distinguishing recovery objectives and the need to test recovery.
- OpenTelemetry: TracesOfficial explanation of spans, context and links; diagnostic traces are separate from the durable completion record.
- Google SRE Workbook: Canarying ReleasesPrimary guidance for comparing canary and control, representative workloads, metrics, and gradual rollout.
- AWS: Blue/Green Deployments introductionEnvironment and traffic-switching model; rollout success still depends on application state and compatibility.
Practice marks stay in this browser.