Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Deadline-Aware Batch Inference Platform

By Anup Rai8 min readReviewed September 2026

Interview problem: teams need to classify, enrich or summarize millions of records overnight. Build a platform that validates input, schedules model work within quotas, resumes after failure and produces complete, attributable output manifests.

This is a planning exercise. Volumes, deadlines and prices are assumptions; model quality and throughput must be measured on representative input lengths and output limits.

1. Requirements

Functional requirements

  1. Accept a versioned input manifest, task/model configuration and completion deadline.
  2. Validate records and report per-record errors before or during execution.
  3. Schedule fair, bounded parallel work across tenants and approved model pools.
  4. Retry eligible failures, resume interrupted work and cancel a batch.
  5. Expose progress and return one terminal outcome per input record.
  6. Publish an immutable output manifest with model, prompt and input versions.

Non-functional requirements

  1. Complete 99% of admitted standard batches by their agreed deadline; admission and rejection must be reported separately.
  2. Preserve tenant isolation and input/output retention controls.
  3. Prevent retries from duplicating the authoritative record result or budget settlement.
  4. Bound cost exposure and queue length.
  5. Sustain recovery after worker loss without restarting the whole dataset.

Assume transformations have no business side effects beyond writing output artifacts. If the model calls payment, email or account-change tools, it is a workflow-automation problem with a separate authorization and idempotency contract.

2. Estimate capacity from the deadline

Assume ten million records in an eight-hour window, mean 800 input and 100 output tokens.

Quantity Calculation Implication
Required record throughput 10M / 28,800 ≈ 347.2 records/s Include validation and settlement overhead
Input processing About 277,778 tokens/s Prefill demand
Output processing About 34,722 tokens/s Decode demand
Measured worker throughput Suppose 25 records/s at the target quality Workload-specific assumption
Minimum worker count ceil(347.2 / 25) = 14 No headroom or failures
At 70% planned utilization ceil(347.2 / (25 × .7)) = 20 Then test failure and long-tail reserve

A batch submitted four hours before the same deadline needs roughly twice the throughput. Admission must consider remaining work, reserved capacity and observed duration distributions. Counting jobs without their token sizes is misleading.

3. Baseline and failure analysis

Architecture / visual model
flowchart LR IN[Input file] --> SCRIPT[One process loops over records] SCRIPT --> MODEL[Approved model endpoint] MODEL --> OUT[Output file]
Read diagram source
flowchart LR
 IN[Input file] --> SCRIPT[One process loops over records]
 SCRIPT --> MODEL[Approved model endpoint]
 MODEL --> OUT[Output file]

This establishes task quality and a per-record cost/latency baseline. It is appropriate for a small experiment.

Failure Repair Benefit Cost or limitation
Crash at record nine million Durable chunk/record checkpoints Resume unfinished work State and reconciliation
One malformed record stops the run Per-record terminal error contract Completeness and diagnosis Consumers must handle partial success
Ten tenants submit together Admission plus weighted fair scheduling Predictable shares and deadlines Some jobs must be delayed or rejected
Provider throttling causes retry storms Shared token/request limiter and jitter Stable throughput Lower instantaneous parallelism
Prompt changes during execution Immutable run configuration Comparable outputs New run required for a change

4. Detailed architecture

Architecture / visual model
flowchart TD C[Batch client] --> API[Identity and admission API] API --> M[(Immutable input manifest)] API --> JOB[(Run configuration and budget reservation)] M --> V[Validate and partition] V --> Q[Deadline and tenant-fair work queues] Q --> W[Leased inference workers] JOB --> W W --> R[Rate-limited approved model routes] R --> W W --> OBJ[(Attempt output objects)] W --> STATE[(Record outcomes and checkpoints)] STATE --> REC[Reconciler and retry scheduler] REC --> Q STATE --> FINAL[Completeness verifier] OBJ --> FINAL FINAL --> PUB[(Published output manifest)] PUB --> C W --> USAGE[Idempotent usage settlement]
Read diagram source
flowchart TD
 C[Batch client] --> API[Identity and admission API]
 API --> M[(Immutable input manifest)]
 API --> JOB[(Run configuration and budget reservation)]
 M --> V[Validate and partition]
 V --> Q[Deadline and tenant-fair work queues]
 Q --> W[Leased inference workers]
 JOB --> W
 W --> R[Rate-limited approved model routes]
 R --> W
 W --> OBJ[(Attempt output objects)]
 W --> STATE[(Record outcomes and checkpoints)]
 STATE --> REC[Reconciler and retry scheduler]
 REC --> Q
 STATE --> FINAL[Completeness verifier]
 OBJ --> FINAL
 FINAL --> PUB[(Published output manifest)]
 PUB --> C
 W --> USAGE[Idempotent usage settlement]

The output manifest is the commit point for consumer-visible batch results. A completed worker does not make the whole batch complete.

5. APIs and storage

POST /batches accepts an input-manifest digest, task configuration, deadline, output schema and idempotency key. GET /batches/{id} reports accepted, queued, running, reconciling, complete, partial, failed or cancelled. GET /batches/{id}/results returns only an authorized published manifest.

Record Key fields Invariant
Run tenant, ID, input digest, model/prompt/runtime versions, deadline Configuration is immutable after acceptance
Work item run ID, record ID, input hash, state, attempt count, lease/fence Stale workers cannot replace a newer terminal result
Result record ID, output hash or error code, validation, attempt reference Exactly one authoritative terminal outcome per record
Manifest run ID, expected IDs/count, result partitions, checksum Completion requires verified coverage, not queue emptiness

Do not key only by array position if the input can be reordered. Use stable record IDs and include source/version identity in the task key. If identical text appears twice with different business IDs, decide whether deduplication is semantically valid before applying it.

6. Execute and reconcile

  1. Authorize the tenant, inspect manifest size and validate its immutable reference.
  2. Estimate token work and reserve an allowed cost/deadline envelope.
  3. Partition into bounded chunks while retaining record-level outcomes.
  4. Workers claim fenced leases and acquire shared provider capacity.
  5. Persist attempt outputs before conditionally committing the authoritative record result.
  6. On uncertain failure, inspect durable outcomes before retrying; preserve attempt usage even if its output is superseded.
  7. A reconciler finds missing, stuck and retry-eligible records. Permanent validation failures become terminal errors.
  8. Publish only after every expected record has a terminal outcome and the manifest passes integrity checks.

When a queue reports zero messages, in-flight or lost work may still exist. Completeness is a comparison between expected input IDs and terminal outcomes.

7. Deadlines, fairness and failure handling

Event Response
Worker lost after model call Retry after lease expiry within budget; deduplicate authoritative results
Provider outage Eligible fallback or pause; recompute deadline feasibility
Ten percent of records much longer Split workload classes and update remaining-work estimates
Tenant cancels Stop new claims; fence publication according to the cancellation contract
Output schema invalid Bounded repair/retry or terminal validation error, never silent truncation
Deadline no longer feasible Notify status explicitly; do not label the batch on time by excluding unfinished records

Use earliest-deadline prioritization within tenant fairness constraints. Reserving all capacity for the newest urgent job can starve already admitted work. Keep interactive model traffic in a separate protected pool unless a tested scheduler can honor both contracts.

8. Quality, cost and rollout

A hypothetical rate of $1/M input and $5/M output yields $8,000 + $5,000 = $13,000 for ten million first attempts. A 5% retry fraction at the same average token cost adds about $650. Add validation, object storage, transfer, scheduling and operator cost. A provider's batch discount is a current commercial term to verify, not an architectural guarantee.

Choice Benefit Tradeoff
Larger chunks Less coordination overhead Coarser replay and skew
Record-level outcomes Fine-grained recovery and audit More metadata operations
Preemptible compute Lower suitable compute cost Checkpointing and deadline risk
Strict schema validation Reliable downstream ingestion More rejected/repaired outputs

Pilot a small representative dataset, including long inputs and invalid records. Compare quality against reviewed labels and a deterministic baseline where possible. Load-test the entire eight-hour window, inject worker/provider failures, and validate every output ID before expanding.

9. Interview questions and closing

Q1: Can you guarantee exactly one model invocation per record?

Sample answer: Usually not across a timeout with an unknown upstream outcome. I can make the authoritative result and settlement idempotent, preserve attempts and bound retries. Duplicate compute is different from duplicate consumer-visible records.

Q2: How do you prove a batch is complete?

Sample answer: Compare the immutable input's expected record IDs with durable terminal outcomes, verify result object checksums and publish a manifest only when coverage matches. An empty queue or 100% worker progress is insufficient.

Q3: What changes if the deadline halves?

Sample answer: Recalculate required throughput and check provider limits, prefill/decode capacity, I/O and budget. More workers help only if those dependencies scale. Admission must reject an infeasible promise instead of silently missing the deadline.

I would build a versioned, deadline-aware workflow with record-level recovery and a verified output manifest. The main tradeoff is utilization versus the capacity reserve needed for skew and failures.

Recall: Immutable input → Feasible admission → Fair scheduling → Durable outcomes → Verified completeness.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an AI Gateway and Model-Routing Service
NEXT LESSONDesign an Image and Video Generation Platform →

Explore the diagram