Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Content-Moderation Platform

By Anup Rai13 min readReviewed September 2026

Interview problem: moderate text and media at social-platform scale while limiting harmful exposure, avoiding wrongful restrictions, and providing timely human review and appeals.

Moderation applies a platform's policy to content and behavior. A classifier estimates signals; a separate decision process chooses an action such as allowing, labeling, limiting distribution, age-gating, restricting or escalating. Similar words can appear in harassment, journalism or a quotation, so a keyword is not a complete policy judgment.

This is a hypothetical Learnastra interview scenario. Volumes, routing fractions, costs and objectives are assumptions to validate, not measured production results.

1. Define policy and the visible product behavior

Clarify media types, languages, live versus uploaded content, geographic scope, user age restrictions, and what remains visible while assessment is pending. Distinguish immediate containment from completing a review and from any external reporting process.

Functional requirements

  1. Accept versioned posts and media, validate formats, and create a durable assessment job.
  2. Apply approved known-media signals, specialized classifiers and contextual review where needed.
  3. Make policy-versioned decisions with evidence and a reason code.
  4. Enforce visibility/account actions idempotently and prevent stale decisions from overwriting newer ones.
  5. Route urgent risks directly to trained specialists from any stage.
  6. Support human decisions, user notifications and eligible appeals.
  7. Record changes, reversals and evidence for audit; curate feedback for evaluation and future training.

Nonfunctional requirements

  1. Plan for 10M posts/day and 50M daily active users; size media processing from bytes, duration and sampling, not just post count.
  2. Define a bounded initial-visibility deadline, such as an illustrative p95 300ms for ordinary text. Deeper assessment may be asynchronous under the agreed policy.
  3. Target critical review completion within 15 minutes, high-priority within one hour, standard within 24 hours and appeals within seven days, each at a stated percentile. These are scenario SLOs, not universal legal deadlines.
  4. Measure harmful exposure as a fraction of views, alongside category-specific precision, recall and false-positive rates.
  5. Protect restricted media, reporter identities, account data and reviewer access; minimize ordinary logs.
  6. Bound retries, model spend and queues, with reserved specialist capacity and defined outage behavior.
  7. Version content, policy, models and decisions so an appeal or incident can reconstruct what happened.

An exposure objective below 0.1%, recall above 99% and precision above 95% are proposed goals. Evaluate whether they are simultaneously feasible for each category and population; a single aggregate score cannot establish that.

Content categories and action policy

Category Initial policy decision Further handling
Suspected child sexual abuse material (CSAM) Restrict access and route to the approved child-safety process Trained specialists and reporting/evidence procedures
Violence/gore Assess severity, context and imminent danger Imminent risk may require urgent containment
Hate speech Assess protected target, context and threat Serious cases to appropriate reviewers
Harassment Examine targeting, repetition and threats Combine content and behavior evidence
Spam Evaluate content, account and network signals Rate limits, distribution controls or review
Misinformation Apply topic-specific policy and evidence checking Toxicity is not a truth detector
Adult content Apply the platform's age/access policy Suspected child exploitation takes the specialist route

In the US, NCMEC's CyberTipline is a reporting channel for suspected online child exploitation. The service's child-safety/legal owners define applicable reporting and retention procedures. A model score does not determine legal obligations.

2. Work the confusion matrix before selecting thresholds

A false positive flags acceptable content; a false negative misses a violation. Precision is TP / (TP + FP), recall is TP / (TP + FN), and false-positive rate is FP / (FP + TN).

For 10,000 posts with 1% violations, 90% recall and a 1% false-positive rate:

Actual class Flagged Not flagged Total
Violation 90 true positives 10 false negatives 100
Acceptable 99 false positives 9,801 true negatives 9,900
Total 189 9,811 10,000

Precision is 90 / 189 ≈ 47.6%, despite the apparently small 1% false-positive rate. 1 − precision is the fraction of flags that are false; it is not the false-positive rate.

A post-level miss and a view-level exposure have different consequences. One missed viral post can account for many harmful views. Estimate exposure from appropriately sampled impressions with reviewed labels, while protecting user privacy and accounting for sampling weights. Appeals and flagged-only review cannot estimate all missed violations.

3. Baseline and the failures that justify more tiers

Start with durable intake, a policy classifier, a versioned decision store and a staffed exception queue. Decide whether content is held or temporarily visible by risk category. A model timeout is an incomplete assessment, not an “allow” result.

Baseline limitation Added mechanism Benefit Cost or new failure
Repeated known harmful media Approved known-media matching Fast recognition of known items Does not detect every novel item; match governance still matters
Rare contextual cases Bounded contextual model review More evidence-sensitive classification Latency, cost, prompt injection and reviewer errors
Video/audio blind spots Media-specific analysis and coverage records Detects signals absent from text Sampling gaps, transcription/OCR errors and compute
High-risk case waits behind ordinary traffic Direct specialist path and reserved capacity Faster urgent handling Capacity and staffing requirements
Wrongful restriction persists Independent appeal/review workflow Correction and accountability Selected feedback and additional load
Retried old decision undoes an appeal Version-checked enforcement Preserves the current decision Coordination with feed/cache services

PhotoDNA compares image signatures against known-image signatures. It is not facial recognition or a general detector of all new abusive content. Match results require the approved operational process; do not claim zero false positives or infer universal certainty from a hash hit.

4. Detailed architecture and cascade arithmetic

Average ingestion is 10M / 86,400 ≈ 116 posts/second; an assumed 10× peak is about 1,160/s. A single video can require much more processing than a short text post, so provision separate bounded media workers.

For one planning scenario, 5%, 85%, 8% and 2% of all submitted posts are resolved at tiers 1–4. These are routing assumptions, not conditional percentages or observed quality results.

Tier Work Items processed/day Resolved/day
1 Known-media and inexpensive signals 10,000,000 500,000
2 Specialized text/media classifiers 9,500,000 8,500,000
3 Contextual model review 1,000,000 800,000
4 Human review 200,000 200,000

Every item reaching tier 3 has already incurred tiers 1 and 2. Urgent bypasses, appeals, rechecks after edits and representative quality audits add work beyond this simplified base flow.

Architecture / visual model
flowchart TD UP[Versioned content intake] --> STORE[(Restricted content store)] UP --> JOB[Durable jobs and admission controls] JOB --> T1[Tier 1 known-media and fast signals] T1 -->|Needs classification| T2[Tier 2 specialized classifiers] T2 -->|Needs context| T3[Tier 3 bounded contextual review] T3 -->|Unresolved| HUMAN[Human queues by severity and expertise] T1 --> POLICY[Versioned policy decision service] T2 --> POLICY T3 --> POLICY JOB -->|Urgent signal| SPECIAL[Immediate containment and specialist queue] T1 -->|Urgent signal| SPECIAL T2 -->|Urgent signal| SPECIAL T3 -->|Urgent signal| SPECIAL SPECIAL --> HUMAN HUMAN --> POLICY POLICY --> DB[(Decision ledger and transactional outbox)] DB --> ENF[Version-checked enforcement worker] ENF --> FEED[Feed, search and media visibility controls] DB --> NOTICE[Reason notice and appeal entry point] NOTICE --> APPEAL[Appeal review with current evidence] APPEAL --> HUMAN DB -. selected cases and random audits .-> EVAL[Evaluation and curated labels]
Read diagram source
flowchart TD
    UP[Versioned content intake] --> STORE[(Restricted content store)]
    UP --> JOB[Durable jobs and admission controls]
    JOB --> T1[Tier 1 known-media and fast signals]
    T1 -->|Needs classification| T2[Tier 2 specialized classifiers]
    T2 -->|Needs context| T3[Tier 3 bounded contextual review]
    T3 -->|Unresolved| HUMAN[Human queues by severity and expertise]
    T1 --> POLICY[Versioned policy decision service]
    T2 --> POLICY
    T3 --> POLICY
    JOB -->|Urgent signal| SPECIAL[Immediate containment and specialist queue]
    T1 -->|Urgent signal| SPECIAL
    T2 -->|Urgent signal| SPECIAL
    T3 -->|Urgent signal| SPECIAL
    SPECIAL --> HUMAN
    HUMAN --> POLICY
    POLICY --> DB[(Decision ledger and transactional outbox)]
    DB --> ENF[Version-checked enforcement worker]
    ENF --> FEED[Feed, search and media visibility controls]
    DB --> NOTICE[Reason notice and appeal entry point]
    NOTICE --> APPEAL[Appeal review with current evidence]
    APPEAL --> HUMAN
    DB -. selected cases and random audits .-> EVAL[Evaluation and curated labels]

The policy service combines signals and explicit rules; it is not just a majority vote. Keep a decision's classification, temporary visibility and final enforcement status separate. A failed enforcement write must not appear as a successfully removed post.

5. Define the contracts and evidence

POST /content/{id}/assessments
{content_version, submission_id} → assessment_id

POST /decisions/{id}/appeals
{appeal_request_id, reason, additional_evidence_ref?} → appeal_id

GET /content/{id}/moderation-status
→ content_version, current_decision_id, visibility, reason_code, appeal_status

All endpoints enforce the caller's role and object access. External reporters do not gain access to restricted content merely by naming its ID.

Record Essential fields
Content version Immutable media/text reference, hash, author, creation/edit time
Assessment Content version, policy/model versions, category signals, coverage and incomplete checks
Decision Stable ID, content version, decision sequence, policy rule, evidence, actor and time
Action job Decision ID, destination, expected/current version and delivery status
Review item Lane, original escalation time, deadline, language/expertise and lease
Appeal Challenged decision, submitted evidence, reviewer and outcome

An edit creates a new content version and assessment. A verdict for version 4 cannot automatically certify version 5. Retention follows the service's defined legal and operational policy; a database version history is not a reason to keep sensitive media forever.

6. What each automated tier can actually decide

Tier 1: fast signals

Known-media matching, account limits and pattern rules identify candidates cheaply. Keywords can indicate a need for context rather than a violation. Keep original and derived representations linked; normalization is evidence processing, not permission to destroy distinctions in legitimate language.

Tier 2: specialist classifiers and modality coverage

Use detectors evaluated for the platform's categories and media. As a current example, omni-moderation-latest accepts text/images and the standalone endpoint is free, but it does not assess audio. Several categories—including hate, harassment and sexual/minors—are text-only. Unsupported image categories can return zero; that means unassessed, not safe. Inspect applied input types. Do not send known or suspected CSAM to this API; use the dedicated child-safety process. OpenAI moderation documentation.

A multimodal general-purpose model may help with memes and context. OCR, speech transcription and specialist vision models can still be useful. Sampling ten video frames does not prove that every frame is safe. Record which media segments and modalities were actually checked.

Tier 3: contextual review

Supply the applicable trusted policy, submitted content, relevant conversation context and prior signals. Request a bounded schema with violation, no_violation or uncertain, a policy-rule reference and evidence locations. Keep the model's short evidence explanation distinct from a guarantee of correctness.

Schema validity establishes structure, not truth. Refusal, timeout, missing output or unsupported modality leaves the assessment incomplete. A provider's safety refusal is not itself proof of a violation under the platform's different policy. Rolling model aliases can change score distributions; recalibrate thresholds against held-out labels after changes.

Perspective API's official notice says service ends after December 31, 2026, with new usage/quota requests closed after February 2026. It is a migration concern for existing integrations, not a suitable new long-term dependency.

Executable example: missing coverage cannot become an allow decision

This small gate consumes trusted coverage metadata and an evaluated proposal. It demonstrates routing; it does not implement the classifiers or the full policy engine.

def assessment_route(*, urgent, required_checks, completed_checks, proposal):
    if urgent:
        return "specialist_containment"
    if not required_checks or not set(required_checks) <= set(completed_checks):
        return "incomplete_review"
    if proposal not in {"violation", "no_violation", "uncertain"}:
        return "incomplete_review"
    if proposal == "uncertain":
        return "human_review"
    return "policy_decision"  # Neither the detector nor this gate applies an action.

assert assessment_route(urgent=True, required_checks={"text"},
    completed_checks=set(), proposal=None) == "specialist_containment"
assert assessment_route(urgent=False, required_checks={"text", "image"},
    completed_checks={"text"}, proposal="no_violation") == "incomplete_review"
assert assessment_route(urgent=False, required_checks={"text"},
    completed_checks={"text"}, proposal="uncertain") == "human_review"

Coverage should name concrete checks, such as image-violence or text-harassment, rather than only the coarse modality names used in this toy example. A completed check also needs a valid result for the required policy/model version.

7. Human review and appeal lifecycle

The completion clock starts at first qualifying escalation and includes queue waiting and active review. Retries do not restart it. An appeal gets its own clock from submission. Immediate containment protects distribution while those later decisions remain pending.

Architecture / visual model
stateDiagram-v2 [*] --> Submitted Submitted --> Tier1 Submitted --> CriticalQueue: urgent risk Tier1 --> Restricted: approved known-match action Tier1 --> Tier2: needs classification Tier1 --> CriticalQueue: urgent risk Tier2 --> Decided: policy decision Tier2 --> Tier3: needs context Tier2 --> CriticalQueue: urgent risk Tier3 --> Decided: policy decision Tier3 --> CriticalQueue: urgent risk Tier3 --> HighQueue: serious unresolved case Tier3 --> StandardQueue: other unresolved case CriticalQueue --> HumanReview: specialist starts HighQueue --> HumanReview: reviewer starts StandardQueue --> HumanReview: reviewer starts HumanReview --> Decided: decision recorded Restricted --> Decided: restriction recorded Decided --> Appealed: eligible appeal submitted Appealed --> AppealQueue AppealQueue --> HumanReview: appeal reviewer starts Decided --> Closed: process complete Closed --> [*]
Read diagram source
stateDiagram-v2
    [*] --> Submitted
    Submitted --> Tier1
    Submitted --> CriticalQueue: urgent risk
    Tier1 --> Restricted: approved known-match action
    Tier1 --> Tier2: needs classification
    Tier1 --> CriticalQueue: urgent risk
    Tier2 --> Decided: policy decision
    Tier2 --> Tier3: needs context
    Tier2 --> CriticalQueue: urgent risk
    Tier3 --> Decided: policy decision
    Tier3 --> CriticalQueue: urgent risk
    Tier3 --> HighQueue: serious unresolved case
    Tier3 --> StandardQueue: other unresolved case
    CriticalQueue --> HumanReview: specialist starts
    HighQueue --> HumanReview: reviewer starts
    StandardQueue --> HumanReview: reviewer starts
    HumanReview --> Decided: decision recorded
    Restricted --> Decided: restriction recorded
    Decided --> Appealed: eligible appeal submitted
    Appealed --> AppealQueue
    AppealQueue --> HumanReview: appeal reviewer starts
    Decided --> Closed: process complete
    Closed --> [*]

Use durable queues with leases/heartbeats, bounded retry and a dead-letter path. Separate severity lanes and reserve critical capacity. Within each lane, consider deadline, reach and calibrated uncertainty, with aging so ordinary cases are not starved. Language expertise and reviewer wellbeing constrain capacity, not merely the number of available workers.

The review UI shows the content version, surrounding context, applicable policy, evidence and decision history. Blur sensitive media by default and restrict access. For selected audits, conceal the model recommendation initially to measure independent judgment and reduce anchoring. Reviewers can disagree, request more evidence or escalate.

Enforcement must not undo a successful appeal

Write the decision and its action job in one database transaction. A worker applies it using the stable decision ID and a compare-and-set condition on the current content/decision version. A duplicate delivery is harmless only if the destination enforces that identity or the update is idempotent.

If decision 12 restores a post after appeal, a delayed worker for decision 11 must not restrict it again. Checking the ledger and then doing an unguarded remote write leaves a race: the destination must enforce the expected version, or all writes must pass through an authoritative serialized executor. Reconcile feed, search and media caches so the decision reaches actual distribution surfaces.

8. Reviewer capacity and economics

At 2% escalation, 10M posts produce 200,000 human cases/day. If 500 moderators each complete an assumed 200 reviews/day, capacity is 100,000/day and the backlog grows by 100,000/day before appeals and audits. The base flow needs 1,000 moderators working that day at that productivity, plus coverage/headroom. Do not assume every sensitive case takes the same time or that increasing throughput preserves review quality.

Cost worksheet with one denominator

The following are hypothetical per-item processing allowances, not model-provider quotes. Every row uses the earlier daily traffic counts.

Component Calculation Cost/day
Fast signals 10M × $0.0001 $1,000
Specialized classification 9.5M × $0.0005 $4,750
Contextual review 1M × $0.004 $4,000
Human review 200,000 × $0.50 $100,000
Subtotal Sum of all processed stages $109,750

That is about $0.010975/post before media extraction, storage, retries, appeals, audits, specialist work and operational overhead. The $0.50 review allowance is not a wage quote. With token-billed models, compute calls × (input_tokens × input_rate + billed_output_tokens × output_rate) / 1M; include reasoning output, images/audio and context tiers under the actual provider contract.

A free detector endpoint still has quotas and integration costs. It reduces spending only where its category coverage and measured quality satisfy the requirement. Reducing escalation from 2% to 1% would halve the base review volume, but is acceptable only if the changed decision policy maintains the required error and exposure outcomes.

9. Robustness, failure policy and learning

Evasion or failure Defense to evaluate Limitation
Character substitution Locale-aware normalization and variant features Can erase legitimate meaning
Text inside images OCR plus visual/contextual assessment OCR mistakes and visual-only signals
Invisible characters Preserve source and inspect normalized views Some characters carry real linguistic meaning
Context manipulation Relevant conversation/behavior context More data and privacy obligations
Encoded content Bounded decoding and resource limits Arbitrary recursive decoding is unsafe/expensive
Adversarial images Test realistic crops, overlays and compression No finite test proves all future variants are covered
Prompt injection Treat posts/OCR as data; keep policy/action authority external Prompt instructions alone are insufficient
Provider outage Category-specific hold/restrict/review policy Neither blanket allow nor blanket remove fits every risk

Keep derived text, OCR and transcripts linked to their source and coverage. Never place restricted media in ordinary logs or send it to a provider merely because a general multimodal API accepts that file type.

Curate human decisions into datasets after quality checks; do not instantly retrain on every reviewer click. Appeals are selected by who contests a decision. Combine them with representative samples of allowed and restricted content, including weighted sampling when estimating population exposure. Measure reviewer disagreement and policy ambiguity separately from model error.

10. Evaluation and staged rollout

  1. Define category labels, action rules, annotation guidelines and severity before fitting thresholds.
  2. Evaluate by language, dialect, region, media and category, including legitimate contextual uses.
  3. Measure precision/recall/FPR, exposure, appeal reversals and incomplete assessments with clear denominators.
  4. Load-test the initial visibility deadline and review queues under bursts and provider failure.
  5. Test edited posts, duplicate jobs, out-of-order enforcement and appeals racing with old actions.
  6. Shadow new models/policies without changing decisions, then canary an approved scope.
  7. Roll back the model/policy configuration when required, while preserving newer individual appeal outcomes and an audit trail.

Policy owners define permitted interventions; ML owners evaluate detectors; operations owns queue capacity, reviewer quality and wellbeing; platform engineers own enforcement consistency. A dashboard should show oldest queued case and missed deadlines, not just average review time.

Interview follow-ups

1. Why can 99% overall accuracy be useless? With 1% violations, an always-allow classifier reaches 99% accuracy while missing every violation. Inspect the confusion matrix and exposure consequences instead.

2. Why keep the decision service separate from the classifier? A probability/label is evidence. Policy determines the action, context requirements, temporary visibility and appeal process. That separation supports versioning, auditing and different rules without pretending every score is a final verdict.

3. Can a zero category score establish an image is safe? No. The category may not support image inputs, or the model may have missed it. Check modality coverage and the applicable detector's evaluation before interpreting the score.

4. How would you stop stale enforcement after an appeal? Store a new decision version and require the destination or serialized executor to reject writes from older versions. Idempotency alone prevents duplicate actions; it does not order conflicting decisions.

5. Can appeals replace random audits? No. They overrepresent users willing and able to appeal and do not expose all harmful items that were allowed. Use appropriately sampled and reviewed population data as well.

6. What changes when the human queue exceeds capacity? Protect urgent lanes, expose the backlog, adjust staffing/scope and investigate upstream failures. Any threshold change must be evaluated against harm and wrongful-restriction costs; hiding cases from the queue does not solve them.

60-second interview answer

I would separate content signals, policy decisions and visibility enforcement. A bounded cascade handles routine cases, while urgent risks go directly to specialists and incomplete assessments follow an explicit temporary policy. Versioned decisions and guarded writes prevent old jobs from undoing appeals. I would size reviewer capacity from arrivals and handling time, measure exposure and both kinds of error by population, and roll out changes through shadow evaluation and controlled traffic with a correction path.

Remember: Policy → Prevalence → Pipeline → People → Appeals.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an IDE Code Assistant
NEXT LESSONDesign a Fresh Market-Intelligence Search Service →

Explore the diagram