System designby Learnastra

System-design interview · Extended interviews

Design a feature-flag and configuration platform

By Anup Rai

Design how to give selected users a consistent feature version, distribute complete configuration updates, handle disconnected applications and roll back a harmful change.

You will learn to

  • Separate configuration authoring/distribution from application-side evaluation.
  • Calculate a deterministic percentage decision with a stable targeting key.
  • Handle partial rollout, stale configuration, defaults, and rollback under a defined freshness contract.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Caching: cache hits, misses, write policies and invalidation · Replication and durability · Real-time communication: polling, long polling, SSE, and WebSocket · Production readiness: SLI, SLO, observability, and recovery

Workload and timing examples are interview assumptions.

01Problem and scope

A feature-flag platform distributes versioned rules that change application behavior without a binary deployment. A cohort is the group of users or tenants assigned to a feature variant. A snapshot is one complete saved version of its configuration. Stable cohort assignment and complete configuration snapshots let servers reach the same decision from the same inputs. This design normally distributes recommendation changes within ten seconds. A disconnected application may temporarily use its previous configuration, then must disable the recommendation feature when that grace period expires. Flag F7, snapshot C17 and tenant54 illustrate the version race. Feature flags do not replace authorization.

Smallest working design

On one server, start with a configuration file and an if enabled branch. That works until many servers refresh at different times, a customer receives inconsistent decisions, or an operator publishes a malformed rule. The design problem becomes distributing safe versions and defining evaluation behavior, not merely storing booleans.

Clarify the control contract

Candidate: “Does this flag only choose a recommendation experience, or authorize a financial/security action?” Interviewer: “Recommendations; it should normally change within ten seconds.” Candidate: “May a disconnected instance keep the old behavior?” Interviewer: “For a bounded grace period, then fail off.” The answer defines a feasible local-evaluation contract. A disconnected process cannot instantly learn that a central switch changed.

Protocol cases to prove

The example is flag F7, snapshot C17 and tenant54. We will prove stable cohort selection, whole-snapshot installation and what happens when C18 overtakes a slow C17 download. Operators must control the change and see which application instances have adopted it. A percentage slider alone cannot show that.

02Functional requirements

  1. Author typed flags. Support boolean, numeric, string and structured flags in separate environments; save drafts and audit changes without silently overwriting another editor.
  2. Publish validated versions. Validate targeting rules, publish an immutable environment version and support scheduled activation.
  3. Roll back safely. Publish a new version containing prior content; do not move the publication generation backward.
  4. Evaluate locally. Application libraries, called software development kits (SDKs), load a validated snapshot and evaluate flags using trusted user or tenant attributes. Each call supplies a fallback of the flag's expected type, such as false for a boolean, and receives the value plus its reason and configuration version.
  5. Target stable cohorts. Support deterministic ordered rules, percentage rollout and multivariate variants with nonoverlapping ranges.
  6. Show rollout progress. Distinguish draft saved, version committed, distribution announced and instances observed on that version; report bounded telemetry.
  7. Retire flags. Remove obsolete application branches as well as configuration. Deleting configuration alone does not clean up callers that still expect a value.

Evaluation and targeting contract

The control plane edits, validates and publishes; the evaluation path answers application requests. Assume negligible added local latency and normal propagation within ten seconds; these are exercise targets, not automatic SDK guarantees. OpenFeature standardizes provider-facing resolution concepts such as key, default and evaluation context, but does not prescribe hosting or freshness architecture. OpenFeature providers.

Rules specify type conversions and missing-attribute behavior. When a user belongs to multiple tenants, the application verifies which tenant the request acts for and uses that tenant's ID as the targeting key. It must not accept an unchecked tenant ID from request text. Document one stable hashing specification across SDK languages.

Safety and scope

Environment credentials and permissions prevent development edits from becoming production publications. A committed central write does not prove that every production instance has disabled a flag. Flags do not replace security authorization or planning for irreversible database migrations.

03Non-functional requirements

  1. Workload assumption. 20,000 application instances and one billion total flag evaluations/s at peak.
  2. Local evaluation latency. Illustrative p99 below 50 microseconds for bounded rules; measure the actual SDK/runtime before promising this.
  3. Publication and propagation. Validated publication below one second p95; propagation to connected healthy instances within ten seconds p99.
  4. Availability boundary. Last-known-good local snapshots can keep evaluation available while the control plane is unavailable. No separate numeric availability target is assumed here.
  5. Freshness bound for F7. Permit stale use for 60 seconds after the last confirmed configuration-freshness signal; then return typed false with reason stale_config.
  6. Authenticity. A checksum detects corruption; authenticated distribution or signatures with trusted keys establishes authenticity.

Snapshot invariants

Operation Required guarantee
Evaluate related flags One request retains the same immutable snapshot for all related evaluations
Install Never expose a partial version
Finish an old download Never replace a newer committed generation
Roll back Publish a higher generation even when content resembles an earlier release
Repeat an evaluation Same version and context produce deterministic rules

Freshness and security qualifications

Measure elapsed freshness time with the process's monotonic clock, which does not jump when the wall clock is adjusted. Do not trust a caller-supplied time. After restart, require a new confirmation of current configuration before using persisted snapshots. F7's false fallback is a product choice: compatibility flags may need another policy.

Security authorization and irreversible schema transitions remain outside ordinary flags. An urgent kill control that must be authoritative at every operation needs a current server-side gate and its network dependency. Cached flags cannot promise both instant revocation and offline availability.

04Capacity estimates

Assume 20K application instances, 50K flag evaluations/second each during peak, and a 100 KB environment snapshot.

Quantity Calculation Consequence
Evaluation rate 20K × 50K/s = 1B evaluations/s Avoid a network call for each lookup
Full snapshot distribution 20K × 100 KB = 2 GB/update Cache/coalesce global publications
Five-second polling 20K / 5 = 4K checks/s Cheap conditional requests still create load
One cached snapshot/instance 100 KB Small relative to application memory

Interpret the assumptions

These are hypothetical planning values. Streaming notifications can announce a new version while clients fetch it from cacheable storage. Polling remains a recovery path. Batch telemetry or sample it; synchronously logging every evaluation could cost more than evaluating the flag.

Remote-call alternative

A remote RPC for every evaluation would turn 1B/s into an enormous network/control-plane workload. Even a 100-byte request/response envelope is 100 GB/s before transport, and latency would sit on application request paths. Local immutable snapshots instead concentrate work on relatively rare publications and cheap in-process lookups.

Evaluation CPU and telemetry

If the rule evaluator takes two microseconds CPU on average, 1B evaluations/s still consumes 2,000 CPU-seconds/s across the fleet. Complex regexes, unbounded lists or dependency cycles can make local evaluation expensive, so validate complexity and compile rules ahead of use. Cache per-request repeated decisions only when context and snapshot are identical; caching decisions across requests can consume large amounts of memory when many distinct users, tenants or attribute combinations each need their own cache entry.

Publication bandwidth

At ten full publications/hour, 2 GB/update means 20 GB/hour of fleet snapshot delivery before cache reuse and compression. A notification-plus-fetch design lets shared distribution caches absorb this without coupling publication to 20K direct connections. Sampling one in 1,000 evaluations still emits 1M telemetry observations/s; aggregate counts locally and batch exposure records carefully rather than assuming sampling alone makes telemetry free.

05APIs and contracts

Draft editing, publication and evaluation are separate operations. A draft version protects an editor from overwriting someone else’s work; an active generation identifies the configuration committed for distribution. An evaluation returns the generation it actually used, which may lag publication while the SDK fetches and validates the new snapshot.

PUT /projects/shop/environments/prod/draft
{expectedDraftVersion:16,flags:{F7:{type:"boolean",threshold:1000,...}}}
→ {draftVersion:17,validation:"passed"}

POST /projects/shop/environments/prod/publish
{draftVersion:17,expectedActiveGeneration:16}
→ {generation:17,snapshotId:C17,status:"committed"}

getBoolean("F7", false, trustedContext)
→ {value:true,reason:"percentage",variant:"on",generation:17}

Publication outcomes

Draft validation errors identify the offending rule/type/reference without changing active configuration. A failed optimistic version check returns a conflict and the current version for review. Scheduled changes create auditable intents; at activation, the publisher revalidates the expected environment state rather than blindly replaying an obsolete draft.

SDK fallback and request consistency

The SDK API requires a typed fallback and returns diagnostic metadata without throwing ordinary missing-flag/network-init failures into the application path. If SDK hooks modify evaluation inputs, specify when those callbacks run. Also specify which attribute wins when global settings, a client instance and an individual call supply the same name, following the selected API contract. A caller cannot silently request a string flag through a boolean getter.

Distribution and freshness

Distribution endpoints support conditional version requests and immutable snapshot URLs. A stream notification says a newer generation exists; it does not carry an unvalidated partial mutation that must be immediately applied. Acknowledgment telemetry reports installed and observed generations separately from central publication. Protect environment credentials and avoid exposing confidential targeting data to browser clients.

06Data model and access patterns

API/record Example
Draft edit PUT /projects/shop/environments/prod/draft {expectedDraftVersion:16,flags:{F7:{type:"boolean",threshold:1000,...}}}
Published snapshot C17,environment=prod,flags...,checksum,createdBy=the operator
Rule F7,type=boolean,seed=recoA,targetKind=tenant,threshold=1000
Context {targetingKey:tenant54,country:US,plan:team}
Evaluation getBoolean(F7,default=false,context) → value, reason, variant, generation

Stable targeting identity

A targeting key identifies the user or tenant being assigned to a variant; attributes such as country or plan supply additional rule inputs. Document their types and which source wins when the same attribute appears more than once. OpenFeature evaluation context. Validate references/types/cycles and rule complexity before publication. Use optimistic version checks to prevent one operator overwriting another’s newer edit.

Publication authority

Keep draft versions, validated snapshots, an active-generation pointer and audit records. A snapshot includes environment, generation, schema/compiler version, canonical content hash and rule dependency metadata. Publishing commits the active pointer and audit record only after the immutable snapshot is durably available. A distributor can replay publication events if notification fails.

SDK-local state

The SDK stores one active immutable snapshot pointer plus initialization/freshness metadata. A background worker compiles the new rules while requests use the old snapshot, then switches the active pointer. A request captures that pointer once if F7 depends on another flag; keeping that reference is called pinning the snapshot. Both evaluations then use the same version, rather than combining values the operator never published together. Previous snapshots can be retained briefly for active requests and debugging, then reclaimed when no references remain.

Typed trusted context

Targeting context is typed data with a stable targetingKey and explicit attributes. Do not use a raw email as a telemetry identifier when a scoped pseudonymous key suffices. The evaluator's hash specification names encoding, field boundaries, algorithm and unsigned conversion; simple string concatenation without delimiters can produce ambiguous inputs. Our modulo example teaches deterministic cohorting, not a promise of a specific vendor's allocation algorithm.

07Basic working design

Deterministic cohort calculation

In this expression, encodeTuple preserves the boundaries between the seed and targeting key, the hash turns that tuple into a repeatable integer, and mod 10,000 takes the remainder, giving a bucket from 0 through 9,999. The comparison then makes a deterministic decision for that targeting key.

Choose user or tenant assignment

Choose the unit: a user rollout and a tenant rollout differ. All users in tenant54 share a tenant-based decision. A changed seed deliberately reshuffles assignment. Experiments with several variants require separate, nonoverlapping bucket ranges and records of which variant each user actually experienced; do not assume boolean rollout rules fully specify experiment analysis.

Evaluate one immutable snapshot

On one application instance, load C17 from a validated file at startup, compile its ordered rules and keep an immutable pointer. Request req61 provides tenant54 from the authenticated application context. The evaluator checks explicit targeting rules first, then percentage fallback; bucket 731 is below 1,000, so it returns true with generation 17 and reason percentage. The application runs the new recommendation branch.

Publish atomically

The publish process writes a complete new file/snapshot and atomically replaces the active pointer only after parsing and validation succeed. A malformed file does not partially replace F7 while leaving its dependencies from an older configuration. Missing flag/type mismatch returns the caller's typed fallback with a diagnostic reason.

When this baseline is enough

This baseline is a useful production pattern for a small service. It proves deterministic evaluation and safe local replacement. As the application grows, operators need to distribute updates, track concurrent edits and see how old each instance’s configuration is. A control plane supplies those functions while evaluation remains local.

architecture · baselineBaseline: deterministic evaluation on one snapshot

One immutable file and stable targeting key are enough for a correct single-instance rollout.

Baseline: deterministic evaluation on one snapshotOne immutable file and stable targeting key are enough for a correct single-instance rollout. operator to file: Publish complete C17; file to sdk: Validate and atomically install; request to sdk: F7 / tenant54 / default false; sdk to request: Bucket731 <1000; true / C17Publish completeC17Validate and atomically installF7 / tenant54 / default falseBucket731 <1000; true / C17ACTOROperatorSTOREValidated configsnapshotSERVICELocal evaluatorACTORApplication requestsync
Read each connection in order
  1. syncPublish complete C17Operator → Validated config snapshot
  2. syncValidate and atomically installValidated config snapshot → Local evaluator
  3. syncF7 / tenant54 / default falseApplication request → Local evaluator
  4. syncBucket731 <1000; true / C17Local evaluator → Application request

08Find the baseline flaws

Failure test What breaks and what must follow
Unstable cohort assignment Randomly choosing ten percent on every request makes tenant54 alternate between experiences. A session may create data under the new path and read it under the old path moments later. Stable hashing fixes cohort stickiness, but only if every SDK agrees on the targeting unit, seed, encoding and algorithm. Different language defaults can otherwise produce different buckets for the same tenant.
Partially mutated configuration A mutable configuration map creates another failure. An updater changes F7, then its dependency F8, while a request reads between those writes. It observes a combination that the operator never validated. Locking each individual flag does not provide a whole-request snapshot. Build and install the complete immutable version instead.
Out-of-order download Now add network distribution: app8 starts downloading C17; C18 arrives quickly and is installed; the slow C17 fetch finishes afterward. Blind “last download completed wins” rolls the instance backward. A real rollback is a newly published generation containing prior intended values, not an older network response replacing a newer one.
Disconnected instance Finally, a central off switch cannot reach app9 during a network partition. Keeping cached behavior indefinitely violates urgent-disable expectations; failing every evaluation on any network hiccup defeats offline continuity. The design needs a per-flag/environment freshness and fallback contract, displayed in operations.

09Improve the design, step by step

1. Audited versioned control plane

  • Trigger: many operators and environments can overwrite files incorrectly.
  • Mechanism: Draft validation checks types, dependencies, bounds and rule complexity; optimistic publication commits an immutable snapshot plus audit. This makes every change reproducible.
  • Benefit, cost and alternative: Costs are workflow and schema/compiler operations; a bad validator can block safe updates or accept harmful rules. File-based configuration remains simpler for a small trusted deployment.

2. Notification plus cacheable snapshot distribution

  • Trigger: 20K instances polling or downloading on every request.
  • Mechanism: A stream announces generation changes, instances fetch immutable content through regional caches, and periodic conditional polls repair lost notifications. Propagation is efficient and recoverable.
  • Benefit, cost and alternative: This adds persistent connections and caches that may be stale; instances may run different versions during rollout. Polling is still needed because reconnecting streams can miss events. Short polling is adequate when fleet size and freshness requirements are modest.

3. Monotonic atomic SDK installation

  • Trigger: partial maps and out-of-order downloads.
  • Mechanism: Parse and compile in a background worker, outside the application request path, validate environment/schema/authenticity, and atomically install only a generation newer than the current one. Requests pin one snapshot. This prevents backwards or mixed configurations.
  • Benefit, cost and alternative: Costs are temporarily retaining multiple compiled snapshots and careful concurrency handling. A global lock around every evaluation is simpler but can become a latency bottleneck.

4. Bounded freshness and batched exposure telemetry

  • Trigger: disconnected instances and invisible cohort failures.
  • Mechanism: The SDK reports version/age/reasons, applies the documented stale fallback, and aggregates evaluation/exposure metrics. Operators can then measure which instances have installed the rollback and which are still using an older version or fallback.
  • Benefit, cost and alternative: Costs include false fallback during partitions and telemetry overhead. If a sensitive action requires a current central decision, check a remote authority and state what happens when it is unavailable. A local flag cannot supply that guarantee.

10Detailed architecture

Audited publication path

Operators enter an authenticated control API that checks environment permissions and validates drafts. The publisher writes immutable snapshots, commits the active-generation pointer and audit record, then emits a publication event. A distributor sends version announcements and serves snapshots through regional caches. The active pointer records the version the publisher committed. Operators separately measure which application instances have installed and used it.

SDK state and evaluation

Each application SDK has a fetch/validation worker, immutable compiled snapshot and local evaluator. The evaluator does not call the control API per request. The fetch worker verifies the requested environment, content integrity and trusted origin/signature before installation. Polling recovers missed stream notifications. On startup, the SDK uses a valid persisted snapshot within its policy or returns typed fallbacks while initializing.

Request context and telemetry

The request path derives trusted context and pins a snapshot before related evaluations. Exposure telemetry is asynchronous and bounded; a telemetry outage cannot block a recommendation request. A remote evaluation service can support confidential server-only rules or selected sensitive decisions, but that optional path has a different latency/fallback contract.

Rollback is another publication

The final diagram shows rollback returning through the publisher as C18 or later. It never draws an operator reaching into 20K mutable process maps. A disconnected app can only react to information it has or its local freshness deadline; the architecture makes that limitation visible.

architecture · finalFinal: audited publication and local evaluation

Central commit, instance installation and request exposure are distinct observable milestones.

Final: audited publication and local evaluationCentral commit, instance installation and request exposure are distinct observable milestones. operator to control: 1. Edit / rollback with expected version; control to publisher: 2. Validated publish intent; publisher to store: 3. Commit snapshot + active generation; store to events: 4. Durable publication event; events to dist: Announce generation; dist to fetch: 5. Push version hint; fetch to cache: 6. Fetch immutable C17; cache to store: Origin snapshot read on miss; fetch to snapshot: 7. Validate; atomically install if newer; app to eval: 8. Trusted context / typed flag; eval to snapshot: 9. Pin one request snapshot; eval to app: 10. Value / reason / generation; eval to telemetry: 11. Aggregate exposure events; fetch to telemetry: Installed generation and age; fetch to store: Confirm active generation; renew bounded freshness1. Edit / rollback with expectedversion2. Validated publish intent3. Commit snapshot + activegeneration4. Durable publication eventAnnounce generation5. Push version hint6. Fetch immutable C17Origin snapshot read on miss7. Validate; atomically install ifnewer8. Trusted context / typed flag9. Pin one request snapshot10. Value / reason / generation11. Aggregate exposure eventsInstalled generation and ageConfirm active generation;renew bounded freshnessACTOREnvironmentoperatorsG1SERVICEDraft / validation APIG1SERVICEConditional snapshotpublisherG1STORESnapshots / activepointer / auditG1QUEUEPublication eventsG1SERVICEVersion distributorG2CACHERegional immutablesnapshot cacheG2WORKERSDK fetch / validationworkerG3STOREAtomic compiledsnapshot pointerG3SERVICELocal typed evaluatorG3ACTORApplication requestcontextG3STOREBatched exposure /version metricsG4syncasynccontrolG1 Audited control planeG2 Version distributionG3 Application runtimeG4 Asynchronous observation
Read each connection in order
  1. sync1. Edit / rollback with expected versionEnvironment operators → Draft / validation API
  2. sync2. Validated publish intentDraft / validation API → Conditional snapshot publisher
  3. sync3. Commit snapshot + active generationConditional snapshot publisher → Snapshots / active pointer / audit
  4. async4. Durable publication eventSnapshots / active pointer / audit → Publication events
  5. asyncAnnounce generationPublication events → Version distributor
  6. control5. Push version hintVersion distributor → SDK fetch / validation worker
  7. sync6. Fetch immutable C17SDK fetch / validation worker → Regional immutable snapshot cache
  8. syncOrigin snapshot read on missRegional immutable snapshot cache → Snapshots / active pointer / audit
  9. sync7. Validate; atomically install if newerSDK fetch / validation worker → Atomic compiled snapshot pointer
  10. sync8. Trusted context / typed flagApplication request context → Local typed evaluator
  11. sync9. Pin one request snapshotLocal typed evaluator → Atomic compiled snapshot pointer
  12. sync10. Value / reason / generationLocal typed evaluator → Application request context
  13. async11. Aggregate exposure eventsLocal typed evaluator → Batched exposure / version metrics
  14. asyncInstalled generation and ageSDK fetch / validation worker → Batched exposure / version metrics
  15. syncConfirm active generation; renew bounded freshnessSDK fetch / validation worker → Snapshots / active pointer / audit

11Write path and acknowledgement

Validate a complete immutable snapshot before activating its generation. An old download must not overwrite newer installed configuration.

Numbered publication and install trace

  1. Validate the draft. The operator edits F7 against version 16. Validation confirms rule types, rollout bounds, and environment permissions.
  2. Publish and audit. The control plane atomically publishes C17 and its audit record; older snapshots remain available for rollback.
  3. Distribute and atomically install. A distributor announces 17. Instance app8 fetches C17, verifies integrity, and installs the entire immutable snapshot atomically.
  4. Build trusted request context. Request req61 obtains tenant54 from trusted application context and asks for F7 with fallback false.
  5. Evaluate the stable bucket. The evaluator checks ordered targeting rules, then compares bucket 731 with threshold 1,000: true.
  6. Use the variant and record exposure. app8 uses recommendations-v2 and asynchronously records exposure with flag/config/variant identifiers.

One request can pin its snapshot if several related flags must agree. Fetching each flag independently from different versions can expose a combination no operator ever published.

Validate before committing

Before publishing, compile the entire dependency graph and reject cycles, missing references, unsupported types and out-of-range thresholds. Store C17 durably, then conditionally advance the environment's active pointer from 16 to 17 with the audit/publication event in one transaction. If another editor already published 17, the operator receives a conflict and reviews the new state instead of overwriting it.

Handle repeated announcements

The distributor may announce 17 repeatedly. app8 fetches by immutable snapshot identity, validates and compiles it, then compares generation against its current pointer. A duplicate 17 is a no-op. An older 16 is rejected; a valid 18 supersedes 17. Installation switches one pointer to the complete new snapshot; requests never see a partly updated map.

Record exposure separately

After req61 uses generation 17, it records an exposure only if the flag actually influences the experience under the chosen analytics definition. Merely evaluating a flag for debugging or a hidden branch should not automatically count as experiment exposure. If the operator rolls back, the operator publishes a new generation with the intended earlier content and observes version adoption and outcome recovery.

12Read and delivery path

Evaluation uses one whole snapshot and a stable cohort key. Expired stale-use grace applies the declared fallback.

Numbered evaluation flow

  1. Establish trusted targeting context. req61 authenticates tenant54 and builds typed context. Context merge precedence is deterministic; an untrusted request attribute cannot override the trusted targeting identity.
  2. Pin a permitted snapshot. The application pins the current immutable snapshot pointer, C17. It verifies that initialization/freshness policy permits using it; otherwise the result is fallback false with a precise reason.
  3. Evaluate bounded rules. The evaluator locates F7 and checks requested type. It evaluates ordered explicit rules with bounded work, then hashes the documented seed/targeting tuple for percentage allocation.
  4. Resolve all dependent flags on one version. Bucket 731 is compared with threshold 1,000. The result is true, variant on, generation 17. A second dependent flag in this request uses the same pinned snapshot even if C18 installs concurrently.
  5. Run the branch and bound telemetry. The application executes the branch and queues bounded aggregated telemetry. Queue saturation drops or samples diagnostic events according to policy rather than delaying the user.

Pinning, defaults and dependencies

On the next request, the SDK can pin C18. This lets a process change behavior atomically at request boundaries without interrupting an active request halfway through its related decisions. Long-running workflows may need to persist their chosen configuration/version so a later retry does not silently change an already-started irreversible plan; ordinary request-local pinning does not cover that larger lifecycle.

13Correctness deep dive

Monotonic installation rule

install(download):
  verify trusted origin/signature, environment and schema
  verify content hash; parse and compile all rules
  require dependency graph and typed values are valid
  repeat:
    old = atomicLoad(activeSnapshot)
    if old != NONE and download.generation <= old.generation: return STALE_OR_DUPLICATE
    if compareAndSwap(activeSnapshot, old, compiledDownload):
      record installed generation; return INSTALLED

evaluateRequest(context):
  snapshot = atomicLoad(activeSnapshot)
  if snapshot == NONE: return typed fallbacks with reason NOT_READY
  if not freshnessAllows(snapshot.generation): return typed fallbacks with reason STALE_CONFIG
  return evaluate all related flags against snapshot

Competing download outcomes

C18 finishes first: app8 installs 18; delayed 17 compares against 18 and is ignored. C17 finishes first: it installs 17, then 18 replaces it. Both orders end at 18. A request already holding 17 completes consistently under 17, while a later request sees 18. Garbage collection cannot free 17 until those references finish.

Malformed C19: validation fails before the pointer change, so 18 remains active. An authenticity check must be more than a checksum supplied alongside the same untrusted bytes: the SDK must authenticate the distribution server or verify a signature with a signing key it already trusts. Concurrent publishers: the active-pointer compare prevents the operator's stale draft from overwriting a newer committed release without review.

Partition behavior

During a partition, app9 cannot download 18. Its monotonic freshness timer eventually triggers the declared fallback. Version checks prevent older downloads from replacing newer ones; the timer limits stale use. Neither sends an instant instruction to an offline app. Report generation and age so operators can distinguish a published rollback from one adopted by the whole fleet.

Authenticity does not establish freshness

Atomic freshness renewal

Under a short SDK update lock, renew only if the response confirms the installed generation and no higher observed generation supersedes it. Once the SDK learns a newer generation exists, it cannot extend the old generation’s deadline while downloading the replacement. A later confirmation of the same still-current generation can renew freshness without reinstalling the snapshot. Installation and confirmation metadata use the same lock/version guard so a late C17 response cannot renew C18's deadline by accident. Evaluations check the pinned generation and its deadline; on restart, return fallback until current authority is confirmed rather than inventing a new age for persisted bytes.

Install and renewal decision table

Event Install snapshot? Extend stale-use deadline?
Valid newer immutable snapshot arrives Yes, after generation guard No; bytes alone do not establish current authority
Authoritative poll confirms the installed active generation No change needed Yes, from that poll's start time and generation guard
Stream keepalive, stale response, or cached C17 fetch No authority change No
Grace expires or startup has no confirmed age Keep data for recovery Return typed fallback until confirmation
sequence · out-of-orderA slow C17 cannot replace C18

Monotonic generation installation handles both completion orders; active requests keep their pinned snapshot.

A slow C17 cannot replace C18Monotonic generation installation handles both completion orders; active requests keep their pinned snapshot. dist to fetch: Announce C17; slow fetch starts; dist to fetch: Announce C18 rollback generation; fetch to active: Validate C18; atomically install if newer; request to active: Pin generation 18; active to request: Immutable C18 reference; fetch to active: Delayed C17 attempts install; active to fetch: 17 <=18: reject stale generation; request to request: Evaluate all flags using18PARTICIPANTDistributionPARTICIPANTSDK fetch workersPARTICIPANTAtomic snapshotpointerPARTICIPANTApplicationrequest1. Announce C17; slow fetchstarts2. Announce C18 rollbackgeneration3. Validate C18; atomicallyinstall if newer4. Pin generation 185. Immutable C18 reference6. Delayed C17 attemptsinstall7. 17 <=18: reject stalegeneration8. Evaluate all flagsusing18syncreturn
Read each connection in order
  1. syncAnnounce C17; slow fetch startsDistribution → SDK fetch workers
  2. syncAnnounce C18 rollback generationDistribution → SDK fetch workers
  3. syncValidate C18; atomically install if newerSDK fetch workers → Atomic snapshot pointer
  4. syncPin generation 18Application request → Atomic snapshot pointer
  5. returnImmutable C18 referenceAtomic snapshot pointer → Application request
  6. syncDelayed C17 attempts installSDK fetch workers → Atomic snapshot pointer
  7. return17 <=18: reject stale generationAtomic snapshot pointer → SDK fetch workers
  8. syncEvaluate all flags using18Application request → Application request

14Failure and recovery

Failure or condition Surviving state, response and recovery
Rollback reaches only connected instances The operator detects elevated errors and publishes C18 setting F7 off. Connected instances update; disconnected app9 still has C17. A rollback button cannot retroactively change an offline cache. Define whether app9 keeps last-known-good configuration, disables after a freshness deadline, or uses a separate authoritative gate for a truly urgent control.
Unsafe fallback or evaluator Defaults differ by flag: a cosmetic recommendation can fail off; a compatibility mode may require the previous known behavior. On startup without a snapshot, return the typed fallback and a reason, then retry initialization. Test malformed updates without replacing the last usable snapshot. Preserve rollback history so undo is publication of a deliberate version, not an unaudited mutation.
Lost stream or duplicate publication If the stream disconnects but snapshot fetching works, periodic version polling detects the missed generation. If the cache serves an older snapshot for a newer announcement, the SDK rejects the mismatch and retries a bounded fresh path. If every distribution endpoint fails, local evaluation continues only through the declared grace, then returns the appropriate fallback. Freshness signals must be authenticated/version-aware; arbitrary successful HTTP responses cannot extend an old snapshot forever.
No valid snapshot or corrupt cache If an instance starts with no valid snapshot, its fallback reason makes reduced functionality observable without crashing all requests. If a dependency flag is missing or types conflict, validation rejects the snapshot rather than creating runtime surprises. If a running application's code no longer supports an old schema, the publisher's compatibility policy must prevent delivering it; a rollback of values is not necessarily compatible with an irreversible code/data change.
Telemetry outage Avoid a telemetry feedback loop in which a metrics outage causes every evaluation to log a synchronous error. Bound logging and report aggregate failure counts. A rollback drill includes an intentionally disconnected instance and verifies the stale fallback deadline, rather than declaring success from the dashboard's publication response.

15Operations, security, and cost

Writer and distribution security

Authenticate control-plane writers, separate development/production permissions, protect distribution credentials, and audit who changed what. Minimize user attributes in evaluation telemetry. Measure publication-to-instance lag, snapshot-version distribution, default/error rate, evaluation latency, exposure counts, and outcome metrics by cohort; an overall healthy average can hide a broken enabled cohort.

Flag retirement and impact metrics

Schedule flag retirement after rollout stabilizes so abandoned branches do not accumulate permanent configuration complexity. Rehearse two concurrent editors, a rejected snapshot, an offline instance, targeting-key migration, and a rollback during an application deployment. A rollout is safe when its state and failure behavior are explainable, not because its dashboard contains a percentage slider.

Reject malformed publications

A concrete bad update is C19 specifying a percentage outside the supported range. Reject it before publication and keep C18 active; a dashboard success response must mean the validated version was committed. If validation instead succeeds but distribution stalls, show committed version and observed application versions separately. This distinguishes an editing failure from a rollout that has not reached every process.

Measure fleet convergence

Observe the histogram of active generations across the fleet, not only the latest committed version. Measure connected-instance propagation p99 against ten seconds, fallback/error rate, snapshot age, compilation time and per-cohort outcome metrics. A rollout can be fully distributed yet harm the enabled cohort; a safe system links exposure to the actual generation/variant used.

Memory and delivery cost

At 100 KB/snapshot, keeping two compiled generations may still be small compared with application memory, but large targeting lists or segment data can change that. Track compiled size and evaluation work per rule. Publishing ten times/hour transfers roughly 20 GB/hour across 20K instances before distribution caching; regional cache hit rate and delta support can reduce origin traffic, with full snapshots retained as a recovery path.

Cross-language and seed migrations

Before changing a seed or targeting key, test every SDK language against the same fixed inputs and expected bucket results, often called golden test vectors. Tenant54 must map to the same bucket in every supported SDK; changes intentionally reshuffling users require a migration plan and cohort analysis. Scheduled flag retirement prevents permanent branching, abandoned credentials and untested combinations. Remove callers or establish fallback behavior before deleting the flag, and preserve audit history separately from active runtime data.

16Decision ledger and limitations

Local versus remote decisions

Choice Benefit Cost
Local SDK snapshot Low per-request latency; offline continuity Staleness and client rule exposure
Remote evaluation Central rule/attribute control Network latency and service dependency
Edge evaluation Regional latency/cache locality Another distribution tier
Hybrid Select sensitive checks remotely Two failure/default contracts

Client-side confidentiality

Client-side applications must not receive confidential targeting rules or other tenants’ attributes just to evaluate a flag. Use an appropriate server-side boundary. Vendor rollout algorithms are implementation choices; LaunchDarkly documents its own percentage allocation behavior, which need not equal our modulo example. Percentage rollouts. Keep targeting-key type stable through migrations or explicitly measure cohort changes.

Further design choices

Additional decision Benefit Cost/limit Revisit when
Whole immutable snapshots Consistent related evaluations Full download/compile and temporary old copies Very large configs justify validated deltas plus full recovery
Stable cohort hashing Sticky rollouts and monotonic expansion Key/seed changes reshuffle users Product intentionally changes targeting unit
Monotonic publication generations Out-of-order fetch safety Rollback is another publication Never replace with wall-clock arrival order
Bounded stale fallback Defined offline behavior Feature availability drops during partition Different flag semantics require another fallback

Network and freshness cost

Remote evaluation can hide confidential rules and centralize decisions, but it adds a network dependency to every call unless cached, at which point staleness returns. Client-side browser SDKs should receive only rules/data safe for that client to inspect; obfuscation does not make downloaded targeting secrets private. Server-side evaluation may use richer context under a controlled boundary.

Experiments need measurement

Experiments require more than a flag: stable assignment, exposure definition, metrics and statistical analysis. A ten-percent threshold is a rollout mechanism, not proof that results are unbiased. Similarly, a flag cannot reverse an incompatible schema migration or erase data already written by enabled code. Use staged compatible migrations alongside the rollout.

17Interview closing

Rehearse the architecture and contract

“I separate audited configuration publication from local request evaluation. The control plane publishes immutable configuration snapshots after type, dependency and concurrency checks. Instances learn about versions through streaming hints and recovery polling, fetch validated snapshots and atomically install only newer generations. The documented targeting key and seed assign rollout cohorts deterministically, so percentage expansion preserves existing assignments. Each request pins one snapshot for related flags.

Defend the critical boundary

“If a newer generation overtakes an older download, an atomic generation comparison prevents rollback through network reordering. A real rollback is a new generation with prior intended values. Disconnected instances cannot learn it instantly, so the flag has a defined freshness grace and typed fallback, with version spread visible to operators. The costs are bounded staleness, distribution/telemetry work and rule lifecycle complexity. My next tests are cross-language cohort vectors and a rollback with one instance partitioned.”

Answer the follow-up

If the interviewer changes a flag into an authorization or urgent spending control, move that decision to an authoritative operation gate and explain the latency/availability cost. If configuration grows too large for whole snapshots, introduce validated versioned deltas with atomic reconstructed snapshots and a full-fetch recovery path; do not expose partial live maps.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why not choose a random number for every request to get ten percent?

Reveal a model answer

That measures requests rather than stable customers and can make the same customer repeatedly switch behavior. I hash a stable targeting key with a rollout seed into fixed buckets. Tenant54’s bucket 731 stays below the ten-percent threshold 1,000 until we intentionally change the rule.

What the answer must demonstrate: Define the rollout unit, key, and seed.

Applied · Question 2

Why install C17 as one immutable snapshot?

Reveal a model answer

Related flags and rules were validated together. Atomic installation prevents app8 from seeing some version 16 values and some version 17 values that no operator intended as a bundle. A request can also pin one snapshot if it evaluates multiple dependent flags.

What the answer must demonstrate: Local atomicity does not imply global simultaneous rollout.

Applied · Question 3

The operator rolls back, but app9 is offline. Is the flag off everywhere?

Reveal a model answer

No. app9 can retain C17 until it reconnects or its freshness policy expires. I would expose version spread and define last-known-good versus fail-off behavior per flag. A truly authoritative emergency/security decision needs a stronger online gate than an ordinary cached feature flag.

What the answer must demonstrate: State the limit of a cached kill switch.

Foundation · Question 4

Why separate the operator’s console from the evaluation path?

Reveal a model answer

Edits are infrequent, privileged, and require validation/audit; evaluations are frequent and latency-sensitive. Distributing versioned snapshots lets the application evaluate locally without calling the authoring database for every branch. The two paths have different availability and permission requirements.

What the answer must demonstrate: Differentiate configuration authority from request-time resolution.

Follow-up · Question 5

Can a browser receive the complete production targeting configuration?

Reveal a model answer

Only if that data is appropriate to expose. Rules can contain sensitive customer cohorts or attributes. For confidential targeting I evaluate server-side or distribute a reduced public configuration, while still enforcing authorization on protected operations independently of flag values.

What the answer must demonstrate: Do not turn rollout metadata into an access-control system.

Follow-up · Question 6

Two operators edit version 16 at once. Which change wins?

Reveal a model answer

I require an expected-version check. The first publication creates 17; the second receives a conflict and must review/reapply its change against the new state rather than silently overwriting it. Audit history records both the successful version and any later deliberate rollback.

What the answer must demonstrate: Rollback history and flag lifecycle are product features.

Applied · Question 7

C18 is installed before a slow C17 download completes. What exact mechanism prevents regression?

Reveal a model answer

The SDK validates and compiles the download while requests keep using the current snapshot. It installs only a newer generation, using compare-and-swap to check and replace the active pointer atomically. Delayed C17 cannot replace C18. A rollback also receives a newer generation even when its values come from an older release.

What the answer must demonstrate: Show the generation comparison and request lifetime.

Follow-up · Question 8

The operator presses off while app9 is offline. When is F7 actually disabled there?

Reveal a model answer

It cannot learn the new central value while disconnected. Under our contract it uses the prior snapshot only until its authenticated freshness grace expires, then returns false with stale_config. Operators see its old generation/age separately from central publication. Only an authoritative, generation-bound confirmation renews that grace; cached bytes and a connected notification socket do not. Its local deadline starts with the confirming request, so network delay cannot extend the stated bound.

What the answer must demonstrate: Do not promise instantaneous remote knowledge during a partition.

Blank-page exercise · 45 minutes

Build the answer yourself

Build the operator’s F7 rollout with stable tenant cohorts. Publish C17, lose connectivity on app9, and roll back with C18 while a second operator edits the same flag.

  • Calculate tenant54’s repeatable bucket decision.
  • Separate authoring, distribution, and evaluation paths.
  • Size snapshot distribution versus remote evaluation.
  • Trace atomic installation and request-pinned configuration.
  • Define defaults, stale-cache rollback limits, and privacy.
  • Handle concurrent edits and flag retirement.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a feature-flag and configuration platformWhat keeps a user’s cohort stable?Recall first, then reveal

A deterministic hash of a stable targeting key and rollout seed, compared with a fixed threshold.

Stable key + stable seed = stable bucket.

Return to lesson
Design a feature-flag and configuration platformWhat is the control plane?Recall first, then reveal

The protected system that edits, validates, versions, and distributes configuration; the application evaluates the resulting snapshot.

Author once, evaluate locally.

Return to lesson
Design a feature-flag and configuration platformWhy is a flag not an access-control rule?Recall first, then reveal

Clients and stale evaluators may retain old flag values; protected actions still require authoritative permissions.

Rollout decides exposure, authorization decides access.

Return to lesson

Final revision

Summary and interview notes

Applications evaluate immutable configurations locally for speed. The publisher audits changes; stable cohort hashing repeats assignments, and atomic installation prevents mixed versions. Current-generation checks and declared fallback values determine what an application does when it loses contact.

Remember these points

  • A percentage rollout is stable only while the targeting unit, key, seed, encoding and hash algorithm stay stable.
  • A rollback publishes a higher generation; old downloads never replace a newer installed snapshot.
  • Snapshot authenticity proves origin and integrity, not that it remains the current configuration.
  • Request-local snapshot pinning prevents mixed flag versions but does not synchronize every process.
  • A disconnected application uses old flags only through the declared grace. An urgent permission or spending check must instead consult current server-side policy before allowing the operation.

Interview tips

  • Calculate one cohort bucket and show both orders of the C17/C18 installation race.
  • Specify exactly which response renews freshness, its deadline origin, and the fallback after expiry.

Important qualifications

  • OpenFeature standardizes evaluation interfaces; these publication, installation and freshness guarantees belong to the platform/provider implementation.
  • Preserving a cohort during threshold growth assumes all higher-priority rules and targeting semantics stay unchanged.

Technical references

Practice marks stay in this browser.