Data engineering for AI builds and operates the pipelines that turn source data into reliable inputs for retrieval, training, prediction and evaluation. It includes ingestion, transformation, validation, lineage and the ongoing handling of changes.
Numerical examples below are illustrative unless explicitly sourced.
Follow one document from source to answer or training example
A policy PDF arrives from a document system. It may contain text, scanned pages, tables, effective dates, and access restrictions. The model does not benefit from that material merely because the file was downloaded. The pipeline must preserve meaning and provenance as it transforms the file into something searchable or trainable.
Ingestion records what arrived and from where. Parsing extracts structure and text. Normalization makes representations consistent, such as line endings, while preserving distinctions that matter. Deduplication identifies repeated material. Lineage records which source and transformation produced a derived item. These names describe different responsibilities; calling the whole process “embedding documents” conceals the stages where information can be lost.
Detect content type and choose a suitable parser. A scanned invoice needs optical character recognition; a table may need its headers and cell relationships preserved. Keep a source reference so a reviewer can compare a suspicious value to the original page. Quarantine parsing failures with an owner rather than quietly indexing empty text.
Understand why cleaning and deduplication require judgment
Repeated headers and navigation can crowd search results and waste training capacity. But repeated legal clauses may be meaningful, and two invoices with the same template may represent different transactions. A normalization rule that removes punctuation can corrupt decimal amounts, product identifiers, or code.
Exact deduplication compares hashes of a defined representation. Near-duplicate methods compare overlapping text patterns; semantic methods compare learned representations. Each broadens the matches and introduces different false-match risks. Define what counts as a duplicate for the downstream task and retain provenance for merged items. A copied policy and its updated exception should not be collapsed merely because most words match.
For RAG, duplicates can fill the result list with repeated evidence instead of covering different required facts. For training, duplication can distort the learning distribution and increase repeated exposure to sensitive material. For evaluation, overlap between training/development data and test cases can make results look better than generalization warrants. Those are related benefits of deduplication, but deduplication cannot prove that an external model never encountered a test item.
Detect and transform sensitive data deliberately
Privacy processing needs both a detector and a decision about what may remain. Pattern matching can find structured values such as many email or account-number formats. Named-entity recognition can help find names and locations whose forms vary. Checksums and domain rules can validate some identifier candidates. Tools such as Presidio combine recognizers and transformations; no detector guarantees that every sensitive value is found.
For a support transcript, transform “Email Alex at alex@example.test about order 8271” into “Email [PERSON_1] at [EMAIL_1] about order [ORDER_1]” when those identifiers are unnecessary for the allowed downstream use. Consistent placeholders can preserve relationships within the record. If the task needs a real order lookup, keep the identifier in the authorized operational system and pass it only to that scoped tool; indiscriminate redaction would break the task.
Masking hides selected characters; redaction removes a value; pseudonymization replaces it with a stable token. A reversible token mapping needs separately restricted storage, retention, and access controls. Pseudonymized text may still be personal or re-identifiable through context. Restrict retained raw sources and propagate deletion to transformed copies and mappings.
Evaluate missed sensitive spans and incorrectly removed useful spans by language, document type, and identifier category. Check meaning after transformation: removing a name must not delete a negation or merge two people's records. Apply minimization before embeddings, training exports, and broad telemetry where required by the data-use contract; scanning only the final answer leaves earlier exposure paths open.
Fork the pipeline according to the consumer
RAG needs searchable chunks, embeddings where used, access metadata, versions, and timely updates and deletions. Fine-tuning needs well-formed examples that represent the desired behavior, valid labels, appropriate permissions for that use, and held-out evaluation data. A document allowed for search is not automatically approved for model training.
Use point-in-time correctness for predictive or historical evaluation: the example must contain only information available at the time of the decision. If tomorrow's resolution is included in today's input, the model may appear excellent because the answer leaked into the features. Split related entities or time periods appropriately, not only random rows.
A batch backfill processes an initial corpus. Incremental processing handles changes afterward, using source events or polling as appropriate. Stable source IDs and versions make retries idempotent: reprocessing version 7 should not create unexplained duplicate chunks. Deletion and permission changes must propagate to indexes, caches, and retained artifacts under the system's policy.
A useful operating contract states freshness targets, allowed sources and purposes, failure handling, ownership, and how completeness is measured. The output is not just “pipeline green”; it is usable, traceable data whose limitations are known.
Build a data contract
A data contract should record:
- Stable source identity and source version.
- Ingestion time and event/effective time, with their distinct meanings.
- Permissions, permitted uses and retention/deletion requirements.
- Parser, normalization and transformation versions.
- Schema, quality status and validation results.
- Downstream lineage and publication version.
- Owner, freshness/completeness targets and repair procedure.
A document's effective date differs from the time the pipeline discovered it. Store both when historical applicability matters.
| Stage | Concrete decision | Failure to test |
|---|---|---|
| Ingest | Connectors, checkpoints, size/type limits | Missed updates, duplicates, malicious files |
| Parse | Preserve text, tables, headings, provenance | Lost units, incorrect OCR, broken reading order |
| Normalize | Standardize encoding without losing meaning | Removing a negation or meaningful layout |
| Deduplicate | Exact and near-duplicate strategy | Merging different policy versions |
| Govern | Permissions, privacy, rights, retention | Data used for an unapproved purpose |
| Validate | Completeness, schema, freshness, quality | Empty documents silently accepted |
| Publish | Versioned index/dataset with rollback | Partially updated corpus exposed |
Deduplication can improve training efficiency and reduce memorized repetition, as shown in Deduplicating Training Data Makes Language Models Better. But similar text is not always redundant: a changed date, amount, or exception can matter. Semantic deduplication needs validation and lineage.
Different consumers need different outputs
RAG needs attributable chunks, search representations, current permissions, and freshness. Fine-tuning needs task-aligned examples, label quality, formatting, rights, and train/validation/test separation. Predictive ML needs feature definitions, point-in-time joins, delayed-label handling, and training-serving consistency. Evaluation needs protected representative cases and explicit expected behavior.
The branches can share parsing and metadata infrastructure; do not assume a cleaned document is automatically suitable training data. Removing all personal data may destroy a legitimate task, while failing to minimize it may violate the intended-use contract. Apply the actual privacy policy and access controls.
Leakage and contamination
For a fraud model predicting at transaction time, a feature computed after the eventual chargeback leaks the answer. Reconstruct what was available at prediction time. For related customers/documents, split at the appropriate entity or time boundary so near-duplicates do not inflate results.
Exact matching, n-gram matching, embedding search, and human inspection can detect different forms of overlap. None proves a proprietary model has never seen an evaluation item. Newly collected private cases reduce some contamination risk; they do not justify an absolute guarantee.
Operate updates and deletions
Use idempotent ingestion keyed by source/version, bounded retries, quarantines for bad records, and replayable backfills. Track change-data-capture lag where relevant. Publish consistent versions or clearly define mixed-version behavior.
Deletion includes source copies, extracted text, vectors, caches, derived artifacts, backups under policy, and any memory containing the data. Removing training examples does not automatically remove their influence from already-trained weights; that needs a separate model-lifecycle decision.
Manager ownership
Assign source owners and downstream consumers a shared contract. Monitor freshness lag, missing/failed records, parser error rates, permission propagation, deletion completion, and downstream quality. A “pipeline succeeded” flag is inadequate if half the tables became unreadable.
Prioritize fixes using user impact. A small number of corrupted invoices can matter more than a large volume of duplicate harmless pages. Budget data labeling and maintenance alongside model work.
Recall questions
“Why is the new model worse only on recent documents?” Check ingestion lag, parser changes, versions, and retrieval before blaming the model.
“Can our own test set guarantee no contamination?” No; control collection and access, detect overlap, and state the remaining uncertainty.
“What does deletion mean?” Trace every derived copy and state the limits for trained model weights.
Practice drawing the update, permission-revocation, and delete paths beside the happy-path ingestion diagram.
Work an exact, lexical, and semantic deduplication cascade
Take four records: A is a return policy, B is a byte-for-byte mirror of A, C changes “30 days” to “14 days,” and D paraphrases A. First retain original bytes and provenance. A content hash identifies A/B as exact duplicates under a stated byte or normalization policy. Keep a canonical record plus all source/version references; do not erase the fact that two sources carried the same text.
For lexical candidates, tokenize into overlapping word shingles. If A has shingle set {a,b,c,d} and another record has {a,b,c,e}, Jaccard similarity is |intersection| / |union| = 3/5 = 0.60. MinHash approximates this similarity by comparing minimum-hash signatures. Locality-sensitive hashing groups signature bands to propose candidate pairs, avoiding an all-pairs comparison; exact verification and a tuned threshold decide what to merge. Bands/rows trade missed matches against candidate volume. SimHash instead produces a compact fingerprint from weighted features and compares Hamming distance; it is a different approximation, not another name for MinHash.
Finally, embeddings can propose semantically close records. SemDeDup studies semantic deduplication using representations and clustering. Similarity is not equivalence: C's changed deadline is critical, even if its vector is almost identical to A's. Preserve policy versions, document type, dates, tenant boundaries, and meaningful field changes before merging. D may be redundant for training diversity but still valuable as a separate authoritative source in retrieval. Thresholds belong to the consumer's contract.
Unicode normalization can change meaning
import unicodedata
assert unicodedata.normalize("NFC", "e\u0301") == "é"
assert unicodedata.normalize("NFC", "x²") == "x²"
assert unicodedata.normalize("NFKC", "x²") == "x2"
NFC composes canonically equivalent forms such as an accented character. NFKC also folds compatibility distinctions, including superscripts. In a mathematical document, x² becoming x2 can destroy meaning. Preserve the source and record the transformation; choose normalization by field rather than applying aggressive cleanup to every document. The Unicode normalization standard explains this distinction.
A filter ablation that changes the decision
Assume a reviewed sample of 1,000 documents, of which 100 are low quality. A heuristic removes 80 bad and 90 good documents: removed-set precision is 80/170 = 47.1%, and bad-document recall is 80%. A classifier removes 85 bad and 20 good: precision is 85/105 = 81.0%, recall is 85%. Combining filters may improve or worsen this; overlapping mistakes mean you cannot add their counts.
Run no-filter, heuristic-only, classifier-only, and combined ablations on the same held-out sample. Inspect which good documents were lost—minority-language text and unusual layouts can be disproportionately removed. Then measure downstream retrieval/training quality and cost, not just removed volume. Quarantine uncertain high-value records for review rather than optimizing a deletion percentage.
Orchestration, distributed compute, and lineage are separate jobs
Read diagram source
flowchart TD
S[Source snapshot and rights] --> O[Orchestrator: dependencies and retries]
O --> C[Compute workers: parse, filter, deduplicate, embed]
C --> V[Validated dataset or index release]
S --> L[Lineage: source IDs and versions]
O --> L
C --> L
V --> L
L --> R[Replay, deletion, and impact investigation]
An orchestrator decides when a task runs and whether prerequisites succeeded. Distributed compute divides large transformations across workers. A lineage catalog records which source and transformation versions produced each artifact. A retryable DAG without lineage cannot explain which released answers depend on a corrupted parser. A lineage catalog without bounded retry and reconciliation does not guarantee the pipeline completed. Name an owner and measurable completeness/freshness contract for each boundary.
Interview questions with developed answers
Q1: Why is deduplication important in an AI data pipeline?
Sample answer: It improves different consumers in different ways. In RAG, repeated chunks can occupy the top results and reduce evidence diversity. In training, duplication can skew the distribution and repeatedly expose the same sensitive examples. Across development and evaluation, overlap can inflate measured performance. I use exact and broader matching methods as appropriate, but inspect false matches and preserve provenance. Similar templates may describe different business events, and a small changed exception can matter greatly. Deduplication is a controlled transformation, not permission to delete everything that looks similar.
Follow-up: Does it guarantee uncontaminated evaluation? No; it reduces detectable overlap within data we can inspect.
Q2: How do you keep evaluation results honest against contamination?
Sample answer: I separate development and held-out data, detect exact and near-duplicate overlap, group related records appropriately, and respect temporal boundaries. I track which cases engineers have used for tuning; repeated use can turn a holdout into development data. I audit source provenance and avoid including future outcomes in inputs. For proprietary pretrained models I cannot prove the absence of prior exposure, so I state that limitation and use fresh, task-relevant cases where feasible. No one checksum or benchmark score settles contamination risk.
Follow-up: Why can a random split be misleading? Related records or later information can appear on both sides.
Q3: How would you debug a RAG answer corrupted by an invoice parser?
Sample answer: I trace the answer back through its chunk and parsed fields to the original document. I inspect whether OCR changed a digit, a table header was separated from a value, or normalization altered a decimal. I correct the parser or routing rule and reprocess affected versions, then evaluate similar layouts. I quarantine unresolved cases rather than treat malformed text as trusted evidence. A larger generation model may not recover information already destroyed upstream, so the repair should occur at the stage that lost the meaning.
Follow-up: What metadata makes this possible? Source ID, page or span references, parser version, and transformation lineage.
Q4: What changes when the same source feeds RAG and fine-tuning?
Sample answer: Some ingestion and quality checks can be shared, but the consumer contracts differ. RAG needs retrieval units, permissions, freshness, and deletion propagation. Fine-tuning needs task examples, reliable labels, distribution control, and permission for training use. I version each derived dataset and its lineage. Access to a document does not automatically grant the right to train on it, and deleting an index entry does not undo a model training run. Those distinctions affect governance and lifecycle design from the start.
Follow-up: Should generated metadata be treated as authoritative? No; preserve its origin and validate consequential uses.
Q5: How do you make ingestion reliable under retries and updates?
Sample answer: I use stable source identities and versions, make transformations repeatable, and upsert derived outputs without creating independent duplicates for the same version. I track completed, failed, and quarantined items and reconcile them against the source inventory. Incremental updates handle changes, deletions, and permissions, while periodic checks catch missed events. I measure freshness and completeness, not only job success. Each failure class has an owner and a replay or repair procedure so one malformed document does not silently disappear.
Follow-up: What if events arrive out of order? Use source version or ordering semantics to avoid replacing newer state with older content.
60-second interview answer
The data pipeline determines what the AI system can know and what we can trust about its outputs. I track provenance, permissions, versions, and intended use from ingestion onward. I preserve document structure, check quality and duplicates, handle updates and deletions, and monitor freshness. Retrieval, training, and evaluation can share infrastructure but require different data contracts. I keep evaluation holdouts separate, use time-appropriate features for predictive models, and make data failures visible instead of silently feeding bad inputs into a stronger model.
Final notes
Recall card: Source → structure → permitted use → quality → versioned publication → update and deletion. Preserve evidence for every transformation; a successful job run is not proof that the resulting data is complete or correct.