System-design interview · Extended interviews
Design event-time click analytics
Count validated event-time clicks with durable input and a transactional duplicate-safe counter, then scale processing and report honestly on late data, hot keys and lag.
You will learn to
- Explain a complete durable queue and transactional-counter baseline.
- Keep logical identity, occurrence time and late-update policy explicit.
- Scale with one recoverable stream/output integration while bounding hot keys and backlog.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Message queues, event logs, delivery guarantees, and backpressure · Data partitioning and sharding · Database indexes: B-trees, composite keys and query access · Replication and durability
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Decide what the number means
A click analytics system turns identified events into counts and rankings. Begin by asking whether the metric is received requests, validated clicks, unique people or billable interactions. These require different identities and validation. We count validated clicks by when they occurred, with minute counts, campaign rollups and hourly top 100 ads. Financial settlement and fraud-model training are separate products.
Click C901 belongs to ad A7 and occurred at 09:00:58, but reaches the collector at 09:02:05. It must revise the 09:00 minute, taking the count from 99 to 100. A repeated delivery of C901 must leave 100 unchanged. A second legitimate click with a different accepted identity may count separately; uniqueness of transport identity does not prove legitimacy.
The complete initial flow is validate → durably append → identify the event-time window → atomically record identity and contribution → expose the count. A single database and processor can implement it. Distribution is justified later by throughput, state size and recovery time, while this meaning of the count remains unchanged.
02Functional requirements
Agree on what the service must do before choosing its components.
- Accept validated clicks. Receive bounded batches with trusted event identity, ad identity and occurrence time; report acceptance after durable storage. Identity distinguishes a retry from another event, not a bot from a person.
- Serve counts and rankings. Return event-time minute counts, campaign rollups and hourly top 100 ads under a named metric and validation policy.
- Explain revisions. Expose metric/policy version, update time, revision, progress and preliminary/final status. Let allowed late events revise their original window.
- Inspect delayed evidence. Retain beyond-policy events for investigation. Financial settlement, fraud-model training, historical audit snapshots and privileged historical rebuilds are separately controlled extensions.
03Non-functional requirements
Use these hypothetical requirements for the worked interview. Confirm the assumptions with the interviewer; the numerical targets require testing and are not measured results. For latency, p95 and p99 mean that 95% and 99% of measured delays, respectively, are no greater than the reported value.
- Workload. Plan for one billion events/day, 200,000/s at peak and 200-byte payloads. Include duplicates, hot ads and catch-up work when testing capacity.
- Latency and availability. Target normal accepted-to-dashboard delay within five seconds, ingestion p95 below 100 ms, narrow query p95 below 200 ms and 99.9% ingestion availability under admitted load.
- Durability. Accepted events must survive one zone failure under the replicated-log policy. Retained input and recovery capacity are finite; stop admitting work before accepted input becomes unrecoverable.
- Counting and finality. One trusted event identity contributes once within the supported replay window. Allow two minutes of ordinary lateness after the window end according to declared event-time progress. Final means closed under that policy, not that no older event can arrive. Reports expose incompleteness rather than treating a missing shard as zero; independent live counters do not form one global snapshot.
- Retention and recovery. Retain raw envelopes for 30 days and online event identities for 24 hours. Never treat an expired identity as a fresh click on sender retry or internal replay; older reconstruction needs a controlled rebuild.
- Security and privacy. Authenticate collectors and campaign queries, validate provenance and timestamp ranges, minimize user identifiers and restrict raw evidence. Define deletion treatment for derived data.
04Make one transactional counter recoverable
Start with a durable queue, one processor and one transactional database. The collector validates a request and acknowledges acceptance only after the queue durably stores it. The processor handles C901 as follows:
- Begin the transaction. Check the trusted scoped identity and payload fingerprint.
- Apply a new contribution atomically. Insert the processed identity and update A7’s 09:00 count from 99 to 100 in that transaction.
- Commit. Make the identity and increment durable together.
- Acknowledge input. Only after commit, acknowledge the input to the queue.
If the transaction commits but the input acknowledgement is lost, the queue can deliver C901 again. The saved identity makes the repeated transaction leave the count unchanged. If the transaction fails, neither the identity nor the increment survives, so replay can apply both. Recording identity and count in separate transactions could instead lose or duplicate a contribution.
Queries read committed counters. A small hourly report sums each ad’s minute counters and sorts the complete totals on that machine. Results remain preliminary until the chosen lateness policy closes their windows. No stream framework is needed to demonstrate this working flow.
The baseline reaches its limit when database write throughput, identity storage or recovery time exceeds its budget. A hot ad can dominate one row while other counters remain idle. Partitioning and batching should address those measured limits while preserving the same transaction boundary.
One local transaction records both identity and contribution.
05Understand event time and report finality
Event time is occurrence time; processing time is when a worker handles the event. A watermark is declared progress through event time used to decide when to emit or close windows. Allow two minutes of ordinary lateness after the window end according to that progress. Final means closed under this policy, not proof that no older event can ever arrive.
Beyond-policy data is retained for investigation or a separately controlled rebuild. An expired online identity must not silently become a fresh click merely because deduplication state was removed. This applies to internal queue replay as well as sender retries.
06Count raw bytes, identity state and recovery work
Assume one billion events/day, a 200,000/s peak and 200-byte payloads. Average rate is about 11,574/s, while peak payload is 40 MB/s. The large peak-to-average difference makes durable buffering useful, but retained input is finite and cannot absorb an unlimited outage.
| State or work | Calculation | Consequence |
|---|---|---|
| Raw input | 1B × 200 B = 200 GB/day | Thirty days is 6 TB before copies |
| One-day identities | 1B × 32 B = 32 GB | Deduplication can exceed counter state |
| Hour of minute counts | 1M ads × 60 × 24 B = 1.44 GB | Dimensions and engine overhead add more |
| Ten-minute peak outage | 200K/s × 600 = 120M events | 24 GB payload awaits processing |
| Catch-up at 300K/s | 120M/(300K − 200K) = 1,200 s | Twenty minutes with continued peak input |
The 32 GB estimate covers only raw identity payload, not database indexes, fingerprints and storage overhead. Larger batches reduce per-transaction work but may increase time before a count becomes visible. Country and device dimensions also multiply distinct counter keys. Bound supported dimensions and benchmark the database with realistic duplicates and skew before choosing batch size or the number of partitions.
07Separate evidence, identity and counters
Click ingestion
POST /clicks
| Input | Meaning or source |
|---|---|
| Event ID | Stable identity of one logical event; C901 in the example. |
| Ad ID | The ad receiving the contribution; A7 in the example. |
| Impression provenance | Evidence used to validate where the click came from. |
| Full occurrence timestamp | When the click occurred, including date/time context. |
| Tenant and collector identity | Derived from verified credentials, not trusted from arbitrary payload fields. |
Acceptance follows durable append; it does not mean the event is billable or already visible. Retry an uncertain response with the same identity and immutable payload. Reject conflicting content under one identity.
Identity and counter keys
| Record | Key | Purpose |
|---|---|---|
| Processed event | (tenant, source, eventId) |
Identify one logical input; retain its payload fingerprint. |
| Counter | (ad, minute, metric, policyVersion) |
Identify the report bucket that input changes. |
Deduplicate by trusted event identity before a changed ad field could send a conflicting retry to another counter.
Stored evidence and results
| Record | Information retained | Why it matters |
|---|---|---|
| Raw envelope | Input, occurrence time and receive time | Supports replay and explains delayed delivery. |
| Processed-event record | Scoped identity and fingerprint within the retry horizon | Recognizes a repeated contribution. |
| Window count | Count and update/revision metadata | Serves a result whose age and changes are inspectable. |
| Metric/validation policy | Versioned campaign and dimension interpretation | Prevents a current campaign lookup from silently reassigning yesterday’s history. |
Count-query response
| Returned information | Meaning |
|---|---|
| Committed counts | Values visible after the counter transaction commits. |
| Metric and policy | The definitions used to interpret those values. |
| Updated-through information | How far processing has progressed. |
| Preliminary/final status | Whether the window is closed under the lateness policy. |
This design promises timely useful reports, not one globally instantaneous snapshot across every independently updating shard. A stronger reproducible cross-shard report is a follow-up with additional coordination and retained snapshot costs.
08Partition independent ads and protect hot ones
First use a replicated durable queue or log so intake can continue through bounded processor downtime. Collectors acknowledge only its configured durable append. Retention and admitted input are finite; expose lag and reduce new admission before an outage can make accepted work unrecoverable.
Partition ordinary processing by ad so related window updates have a clear owner. Different ads can progress independently. Batch database work to amortize transaction overhead, while keeping each event identity and count contribution in the same committed transaction. A retrying batch checks every event identity rather than assuming the entire previous attempt failed.
Measure skew before splitting keys. If A7 receives 40,000 events/s, distribute it across 16 partial counter keys using a stable suffix derived from event identity. Each sees roughly 2,500/s on average. A report combines all required partial counts for A7 before presenting its total. Stable event identity still governs whether a retried click contributes again.
The cost is extra partial state and report work; ordinary ads should retain the simpler path. A counter shard outage can make a report late or incomplete, so disclose that state instead of interpreting a missing partial as zero. More API servers do not repair a hot database row; batching, placement and bounded admission address the actual write work.
09Use one stream framework with a defined output boundary
For the larger workload, Kafka can retain identified input and Flink can provide partitioned processing, event-time progress and recovery checkpoints. A checkpoint records processing state and the input positions from which it can resume. Explain what the counter does when recovery repeats an event; naming the framework does not answer that question.
Choose an idempotent transactional database sink for this design. For each accepted contribution, its transaction inserts the processed-event identity and changes the appropriate counter together. If that identity already exists with the same fingerprint, the sink makes no second change. Several events may share a batch transaction. Record that the input has been processed only after confirming the database transaction committed.
A crash after the counter transaction but before checkpoint/progress completion causes replay. The sink finds the saved identities and does not increment again. A crash before commit leaves neither effect nor identity, so replay applies the work. This is the same baseline proof at a larger processing scale.
Flink’s internal recovery guarantee does not make an arbitrary external write transactional. If we later replace the sink with bulk immutable output or another database, verify its checkpoint-aware or idempotent integration explicitly. Detailed generation-publication protocols are unnecessary to explain this selected sink contract.
Never automatically replay input beyond the retained identity coverage. If recovery needs older input, stop and rebuild the affected counters from retained raw evidence under the controlled rebuild process before resuming. Otherwise a previously committed count could be incremented again after its identity expired. Retention is part of this recovery contract, not just storage cleanup.
The chosen database sink supplies the duplicate-safe output boundary.
Read each connection in order
- syncCommit C901 identity and count 100Stream processor → Transactional counter store
- blockedCrash before saving progressStream processor → Stream processor
- asyncReplay C901 after restartInput progress → Stream processor
- syncFind saved identity; no incrementStream processor → Transactional counter store
- syncAdvance confirmed progressStream processor → Input progress
Collectors acknowledge durable Kafka input. Flink processors commit each event identity with its counter change before advancing progress. Reports combine the required partial counts and expose window completeness; a missing partial is not zero.
Read each connection in order
- syncSubmit clickIdentified click events → Validated collectors
- syncAppend before acceptanceValidated collectors → Retained Kafka input
- asyncPartitioned inputRetained Kafka input → Flink processors
- syncCommit identity and countFlink processors → Counter partitions and IDs
- syncRequest time windowDashboard → Report API
- syncRead all required partialsReport API → Counter partitions and IDs
- returnTotals and completenessReport API → Dashboard
10Explain a late event to a dashboard reader
Late-event timeline
| Observation | Result |
|---|---|
| Watermark passes 09:01 | The 09:00–09:01 window may emit preliminary revision 4. |
| C901 arrives while the watermark is 09:01:30 | Its occurrence timestamp still assigns it to the 09:00 window. |
| Apply C901 within allowed lateness | Publish revision 5 with count 100. |
| Watermark reaches 09:03 | The two-minute allowed-lateness boundary for that window ends. |
With several active inputs, the slowest relevant watermark limits combined progress. Otherwise one delayed partition would be excluded while a result falsely claims completeness. An idle-input policy can permit progress, but data arriving when that input returns still follows the late-data rule. Idleness is not proof that old events no longer exist.
A query authenticates campaign scope and reads compatible committed metric/policy rows. It returns their update time, revision and finality. Paginate a saved report when stable pages matter; otherwise disclose that live counts can change between requests. Do not imply that a cursor alone freezes the dataset.
When the window closes under policy, later evidence enters a correction stream. Do not advance watermarks just to improve a freshness chart. A lagging honest result is different from an apparently final result produced by suppressing inconvenient inputs.
11Build useful reports from complete per-ad totals
The dashboard reads committed minute counters, then groups compatible metric/policy values for campaign summaries. A periodic report job sums an ad’s minute and partial counters for the requested hour before sorting ads for top 100. Keep deterministic tie handling and state the report’s build time and completeness. Those independently changing counters may have been read at different times.
Do not rank partial contributors before summing each ad. If Red has 6 on each of two workers, it totals 12 even though local Blue = 7 and Green = 7 beat it separately. The report needs Red’s complete total first. This simple example is enough to explain why a fast local winner list is not automatically a valid hourly report.
Closed windows give more stable reports, but final still means closed under the stated lateness policy. A missing shard or unresolved input delay cannot be silently counted as zero. Return an older labeled completed report, expose incompleteness or fail the precise report request.
Beyond-policy events remain in a correction queue for investigation or a controlled historical rebuild from retained raw evidence. Keep the main product’s live counting path simple. A separately versioned, reproducible global correction/report system is a worthwhile extension when audit requirements demand it, with its ownership and publication protocol designed explicitly.
12Recover input, control cost and audit revisions
Monitor accepted-to-visible lag, oldest needed input offset, checkpoint duration, watermark age, late-event rates, duplicate conflicts, hot-key skew and correction differences. A high ingestion success rate can coexist with a stalled dashboard. Alert before retention deletes input still needed for recovery, and ensure processors have net capacity to drain backlog.
Authenticate collectors and campaign queries. Validate timestamp ranges and provenance so forged far-future events cannot drive progress. Minimize user identifiers, restrict raw-event access and define deletion treatment for derived datasets. A valid event ID is not an anti-bot mechanism: a malicious client can invent fresh IDs unless an upstream trusted contract limits them.
Test crashes before and after the sink transaction, duplicate collector batches, an idle source returning late data and a counter shard disappearing during report construction. Compare totals with an offline recount of a fixed retained source interval under the same policy.
Costs follow raw retention, identity lifetime, checkpoint state and dimensionality. Seven days of 32-GB/day raw identity payload would be 224 GB before indexes. Limiting live retries and correcting older data in a controlled batch can cost less than retaining every event identity forever.
13Check the design against its requirements
Before closing, check the final design against the agreed requirements. FR means functional requirement and NFR means non-functional requirement; the numbers refer to the lists above. These are proposed validation checks, not test results.
| Requirement | Mechanism in the final design | Validation and remaining limit |
|---|---|---|
| FR 1; NFR 3, 4 | Replicated input and a transaction coupling processed identity with the count protect accepted work. | Crash before/after the sink commit, replay C901 and lose a zone; require one contribution and durable accepted input within the supported model. |
| FR 2, 3; NFR 4 | Event-time windows revise counts; reports combine complete per-ad partials and expose status. | C901 must revise 09:00 from 99 to 100. Test the Red/Blue/Green ranking example and a missing shard; do not claim a global snapshot. |
| NFR 1, 2 | Partitioning, bounded batches, hot-key partials and backpressure address write demand. | Benchmark 200,000 events/s with skew and duplicates, measure all latency targets and monitor ingestion availability. Test catch-up while new input continues. |
| FR 4; NFR 5 | Retention checks prevent unprotected replay; beyond-policy evidence enters controlled correction. | Attempt replay after 24 hours of identity coverage has expired; require a stop/rebuild decision, not a second increment. Raw evidence ends at 30 days. |
| NFR 6 | Trusted source identity and campaign authorization restrict input and reports. | Inject forged timestamps, conflicting event payloads and unauthorized campaign queries; check private-data retention. |
14Rapid revision
Rehearse the numbered functional requirements and non-functional targets first. Use this table to recall the mechanisms, then close with the requirements check above.
Remember: Save identity with the count; retries cannot count twice.
| Question | Main-design answer |
|---|---|
| Which time assigns C901 to a window? | Its occurrence time puts it in 09:00; arrival time explains the delay |
| Baseline? | Save the input durably, then save its processed ID and count change in one transaction |
| Why is increment alone insufficient? | Retrying delivery can increment the same event twice |
| What does a checkpoint do? | Save processing state and the input position from which recovery resumes |
| What makes a repeated database update safe? | The processed event identity and its count change commit together |
| What is a watermark? | Reported progress through event time; the lateness policy uses it to close windows |
| How does a hot ad scale? | Split it across stable partial keys, then combine all its partial counts before reporting |
| What happens under overload? | Keep accepted input, show processing delay and slow or reject new arrivals |
| What is deferred? | Reports from one exact global instant and a complete process for correcting historical results |
Close with: “I define validated event-time clicks, then start with a durable log and a transactional counter that records each processed identity. At scale I partition and batch that same work, using one stream framework with a verified idempotent sink. Late data updates its original window under a clear policy. Reports combine complete per-ad totals and honestly show age or incompleteness.”
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What metric are you actually counting?
Reveal a model answer
Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record. Transport deduplication does not establish that a distinct event is legitimate or billable.
Interviewer follow-up
Why keep raw input?
Reveal the follow-up answer
It explains revisions and supports bounded investigation or a controlled recount.
What the answer must demonstrate: Validated identified clicks by occurrence time. Received requests, unique people and billable interactions have different rules, so the policy version and trusted provenance are part of the record.
What is the smallest complete design?
Reveal a model answer
A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter. Queries read committed counts. Saving both effects together makes a lost worker response recoverable without assuming every delivery occurs once.
Interviewer follow-up
Why not expose a stateless increment endpoint?
Reveal the follow-up answer
A timed-out caller can repeat the increment even when the first one committed.
What the answer must demonstrate: A collector durably appends identified events; one worker transactionally inserts the processed identity and increments the correct ad/minute counter.
Where does C901 belong when it arrives at 09:02:05?
Reveal a model answer
Its occurrence at 09:00:58 assigns it to the 09:00 minute. If it is within the configured late-update policy, revise that original window from 99 to 100. Otherwise retain it for the correction process. Processing time must not silently answer a different business question.
Interviewer follow-up
What does a watermark mean?
Reveal the follow-up answer
Declared progress through event time, used to close windows under explicit assumptions; it does not rule out all late events.
What the answer must demonstrate: Its occurrence at 09:00:58 assigns it to the 09:00 minute.
The sink commits and the processor crashes before progress is saved. What happens?
Reveal a model answer
Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change. If the previous transaction had aborted, neither identity nor effect would exist and replay would apply it. This contract must be verified for the actual external store.
Interviewer follow-up
Does a framework checkpoint automatically cover every database?
Reveal the follow-up answer
No. Use a documented transactional/checkpoint integration or a deliberately idempotent sink.
What the answer must demonstrate: Recovery replays the event. The sink transaction finds its saved scoped identity and makes no second count change.
How would you handle one ad with 40,000 events/s?
Reveal a model answer
First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total. Stable event identity still prevents duplicate contribution. Splitting every ordinary key adds unnecessary report and state cost.
Interviewer follow-up
What if a partial is unavailable?
Reveal the follow-up answer
Disclose incompleteness, use a labeled completed report or fail; do not treat missing data as zero.
What the answer must demonstrate: First measure batching and database limits. If one counter remains hot, use stable partial keys, such as 16 salts averaging 2,500/s, then combine all partials before showing that ad’s total.
Why can worker-local top lists miss an hourly winner?
Reveal a model answer
Red = 6 on two workers totals 12, but each worker can select a different local item at 7. A report must aggregate each ad’s complete total before sorting winners. This design uses a periodic report with stated build time; exact globally simultaneous rankings are a stronger follow-up.
Interviewer follow-up
Does a cursor create a fixed report?
Reveal the follow-up answer
No. Stable pagination needs a retained report or snapshot, not just a remembered key.
What the answer must demonstrate: Red = 6 on two workers totals 12, but each worker can select a different local item at 7.
How long does a ten-minute peak backlog take to recover?
Reveal a model answer
There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds. Preserve admitted work, expose lag and tighten intake before retention is exhausted.
Interviewer follow-up
Can progress be advanced to make the dashboard look healthy?
Reveal the follow-up answer
No. That can falsely finalize incomplete windows.
What the answer must demonstrate: There are 120 million events. If processing resumes at 300,000/s while 200,000/s still arrive, spare throughput is 100,000/s and recovery takes 1,200 seconds.
What happens to a very old resend after identity expiry?
Reveal a model answer
A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe. Enforce an admission policy for older work and use controlled reconstruction from retained evidence for historical repair. Do not silently treat forgotten identities as new clicks. Internal recovery obeys the same limit: stop automatic replay beyond retained identity coverage and rebuild affected counters from raw evidence instead of applying old inputs to existing counts.
Interviewer follow-up
When would you add exact global correction publication?
Reveal the follow-up answer
When reproducible audit requirements justify the additional snapshot, ownership and publication protocol.
What the answer must demonstrate: A24-hour online identity window cannot promise that an arbitrary historical resend is duplicate-safe.
Blank-page exercise · 45 minutes
Build the answer yourself
Count late C901 exactly once in the database, replay it after a crash, then scale to a hot ad and a ten-minute backlog.
- Agree the numbered functional requirements and non-functional targets: the counted metric, identity, event time, lateness, load and durability.
- Draw the complete baseline after agreeing the requirements.
- Calculate retained state and catch-up capacity.
- Prove the transactional sink.
- Explain framework recovery and late updates.
- Check counts, report completeness, latency, retention and recovery against the numbered requirements; identify remaining benchmarks and advanced limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design event-time click analyticsC901 commits, but its queue acknowledgment is lost. Why does replay leave the count at 100?Recall first, then reveal
The processed-event record and count change committed together. Replay finds that record and skips the increment; if the transaction failed, neither change survived.
Save identity with the count; retries cannot count twice.
Return to lessonDesign event-time click analyticsA click arrives late. Which timestamp chooses its minute?Recall first, then reveal
When it occurred. Arrival and processing timestamps explain the delay.
Count when it happened.
Return to lessonDesign event-time click analyticsAn ad’s counts are split across workers. What must happen before ranking ads?Recall first, then reveal
Combine all partial counts for each ad, then compare the complete totals.
Sum before top.
Return to lessonFinal revision
Summary and interview notes
Save each event identity with its count change so a retry cannot count it twice. Put allowed late clicks in their original window; route later evidence to corrections. Show report age and completion status so later corrections are understandable.
Remember these points
- Agree the numbered functional requirements and non-functional targets before designing components; validate the final design against them.
- Then establish a simple recoverable counter.
- Use scoped stable event identities.
- Assign windows by occurrence time.
- Verify the framework’s actual output integration.
- Combine hot-key partials before reporting.
Interview tips
- Use a lost sink acknowledgment to explain replay.
- Calculate recovery with spare throughput.
Important qualifications
- Exact global snapshot reports are a separate extension.
- Historical reconstruction is bounded by raw retention.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Checkpoint-generation publication and stale job epochs
Needed for a different output store that stages bulk results; not required to teach the selected transactional sink.
- Historical interval authority and versioned corrections
Add when audit-grade reproducible backfills become a requirement.
- Exact global top-k publication
Stronger than this main design’s periodic reports with explicit age/completeness.
- Artifact lifecycle guards
A storage-level extension for staged checkpoints and retained output files.
Technical references
- Flink event time and watermarksPrimary explanation of event time, parallel watermarks, late records, and windows.
- Flink checkpointingExplains recoverable stream state and the conditions behind checkpoint-based processing guarantees.
Practice marks stay in this browser.