Synthetic data is data generated or constructed rather than directly collected as observations of the target process. For language-model training it may include generated instructions, answers, preference pairs, conversations, or simulated tool trajectories. Synthetic does not mean anonymous, correct, diverse, or free.
Begin with a coverage gap: a classifier lacks examples of conflicting dates, an assistant mishandles missing evidence, or a tool agent fails after a timeout. The goal is to improve a measured behavior, not maximize the number of generated rows.
Establish the dataset contract
- Identify the target task and important slices using an independent evaluation set.
- Specify the allowed sources, generator, intended use, and retained provenance.
- Define acceptance rules before generation.
- Keep the test set separate from seeds, generation prompts, and filtering feedback.
- Compare models trained with and without the added data under a common budget.
Do not quote undisclosed percentages of synthetic text in frontier-model training mixtures. Public evidence for one recipe does not establish what every provider does.
Design the generation and acceptance pipeline
Read diagram source
flowchart LR
A[Reviewed seeds and coverage gaps] --> B[Generate candidates]
B --> C[Schema and provenance checks]
C --> D[Independent validation]
D --> E[Deduplication and slice balancing]
E --> F[Human audit of accepted samples]
F --> G[Versioned training set]
G --> H[Held-out model evaluation]
H --> A
The last arrow changes future training priorities. It must not feed protected test answers into the next generation prompt.
| Stage | Useful check | What the check does not prove |
|---|---|---|
| Syntax | Schema, required fields, parseable tool arguments | Semantic correctness |
| Source grounding | Answer supported by the supplied document | Source truth, currentness, or permission to use it |
| Execution | Code passes isolated tests | Correctness on all inputs or freedom from malicious behavior |
| Judge model | Rubric-based quality comparison | Independent truth or absence of shared model biases |
| Deduplication | Exact hashes and reviewed near-duplicate clusters | Full semantic diversity |
| Human audit | Expert checks by task and risk slice | Every unsampled row is correct |
Run generated code in a sandbox with resource limits and no production credentials. The data pipeline itself must not become an execution path for untrusted model output.
Expand breadth and difficulty deliberately
Evol-Instruct is a documented approach to expanding instruction data through changes that increase task breadth or complexity. WizardLM is its primary reference. For practical curation, separate the requested skill from incidental wording.
For a document-extraction task, a useful progression is:
- Extract an explicitly labeled invoice date.
- Distinguish invoice date from due date.
- Handle a missing date without inventing one.
- Resolve conflicting dates using a stated precedence rule.
- Return an uncertainty flag when no rule resolves the conflict.
Each level needs a correct target and an evaluation slice. Adding more constraints or elaborate language does not automatically create a valuable hard example. Preserve simple examples so training still represents the real workload.
Critique, revision, and AI feedback
A critique-and-revise loop generates an answer, checks it against explicit principles, and revises it. Accepted revised answers can be used for supervised training. Reinforcement learning from AI feedback (RLAIF) instead uses AI-generated feedback as a training signal, often through preference labels and a reward-learning pipeline. These are related, but not identical operations.
Constitutional AI describes both supervised and reinforcement-learning stages guided by a set of principles. In a product, make the rubric concrete: “cite the supporting policy sentence” can be audited more clearly than “be trustworthy.” Calibrate judge decisions against human-reviewed cases and disagreements.
Match verification to the claim
A numeric answer checker establishes agreement with a known answer under its parsing rules. A proof assistant checks a proof in its formal system, subject to the encoded statement and assumptions. Python execution checks what the program and tests actually do. These are different guarantees.
For retrieval data, generate questions from authorized source passages, then independently check that the answer is supported and the question is answerable. Include deliberately unanswerable cases. Otherwise a model may learn that every retrieval result must contain an answer.
Diversity and recursive training risks
Model collapse refers to degradation that can occur when recursively trained generative models lose information about the original distribution, including its tails. It is not an inevitable property of every synthetic dataset. The recursive-training study motivates checking how generated data is mixed, selected, and reused.
There is no universal “5–20% real data” ratio that prevents all problems. Preserve relevant real or independently reviewed examples where permitted, maintain provenance across generations, and compare slice-level outcomes. A stronger teacher and a checker can still share or introduce systematic errors.
Language translation is not sufficient evidence of cultural or linguistic coverage. Review terminology and task validity with appropriate expertise. Perplexity filters can flag unusual text, but may also reject useful rare-domain or minority-language examples. Treat them as signals to investigate, not universal quality scores.
Work through cost and sampling
Suppose generation of 100,000 candidates costs $300. Automated checks cost $200, and human audit costs $500. If 20,000 examples survive, the observed preparation cost is $1,000 / 20,000 = $0.05 per accepted example, before training and refresh work. Improving acceptance can be more valuable than lowering token price.
A single random sample may miss a rare harmful category. Stratify audits by source, language, task, generator version, and risk. Under an idealized independent random sample, zero observed defects in 300 examples gives an approximate 95% upper defect-rate bound of 3/300 = 1%; it does not establish zero defects, and dependent or biased samples weaken the inference.
Interview practice
- When is synthetic data useful? When it addresses a defined gap and improves protected evaluation compared with a suitable baseline.
- Why keep provenance? To trace errors, rights, generator changes, recursive reuse, and removal requests to affected dataset versions.
- Does code execution verify a solution? Only against the executed tests and environment; it is not a proof over all inputs.
- Why can a stronger judge still fail? It can share generator biases, favor style, or misunderstand the rubric. Calibrate it independently.
- How do you prevent benchmark leakage? Isolate test data from seeds and prompts, deduplicate across splits, and inspect near-duplicate and source overlap.
- What should be optimized economically? Cost per accepted useful example and downstream improvement, including validation and refresh cost—not raw row count.
Recall card and closing
Gap → generation → verification → diversity → measured gain. Close with the independent acceptance checks and the evidence that the resulting training data improves the product. Continue to distillation and RLVR for two ways to use generated supervision.