Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Fine-tuning: change behavior for a measured reason

By Anup Rai5 min readReviewed September 2026

Fine-tuning continues training a pretrained model on additional data to adapt its parameters to a task or domain. Supervised fine-tuning (SFT) learns from labeled examples, commonly instructions paired with desired responses. It can teach knowledge as well as behavior, but it is an unreliable substitute for a maintained source of changing facts.

The Learnastra interview approach is to diagnose the failure first. “We will fine-tune the model” is a proposal whose quality, cost, and maintenance burden still need justification.

When to Fine-Tune

Observed failure First baseline Evidence that could justify training
Answers use outdated policies Retrieval from versioned policy documents Repeated failure to use supplied evidence correctly
Output violates a JSON schema Constrained generation and application validation Valid structure but persistently incorrect field meaning
Responses ignore a repeatable task rule Clear instructions and a few representative examples Recurring behavioral errors across a held-out set
A large model is too expensive Smaller model, shorter prompts, caching Adaptation preserves required quality at lower lifecycle cost
Domain language is poorly understood A capable model with representative context Domain evaluations show a representation or task gap

RAG and fine-tuning can work together. Retrieval supplies evidence; adaptation may improve how the model uses it. Neither replaces authorization or business-rule validation.

Build the supervised learning problem

Instruction tuning is SFT using instructions and target responses across selected tasks. In a causal language model, a common recipe computes cross-entropy on the response tokens while masking the prompt from the loss. The input still contains the prompt; loss masking does not remove its context.

  1. Define the desired response and the conditions for abstaining or asking for clarification.
  2. Collect examples covering routine cases, difficult cases, and important languages or customer groups.
  3. Review targets. A wrong answer in an ordinary SFT target teaches imitation of that answer; it does not act as a negative example.
  4. Separate train, development, and protected test data by entity, source, or time where near duplicates could leak.
  5. Train candidate versions with recorded data, tokenizer, base-model, and configuration revisions.
  6. Select using development results; use the protected test set for the final comparison.

Preference learning handles chosen/rejected comparisons with a different objective. Do not silently substitute its data format for an SFT training recipe.

Architecture / visual model
flowchart LR A[Measured product failure] --> B[Prompt and retrieval baseline] B --> C[Reviewed task examples] C --> D[Train candidate versions] D --> E[Independent task and regression tests] E --> F{Meets release criteria?} F -->|Yes| G[Canary and rollback-ready release] F -->|No| C
Read diagram source
flowchart LR
  A[Measured product failure] --> B[Prompt and retrieval baseline]
  B --> C[Reviewed task examples]
  C --> D[Train candidate versions]
  D --> E[Independent task and regression tests]
  E --> F{Meets release criteria?}
  F -->|Yes| G[Canary and rollback-ready release]
  F -->|No| C

Continued Pretraining

Continued pretraining further trains on text using a language-modeling objective, often to adapt to a domain. It does not require manually written instruction/answer pairs. It is useful to investigate when domain language or representations are the bottleneck, but raw exposure does not directly teach the final application's answer format or tool-use rules.

A document-processing model might first learn from authorized technical manuals, then receive SFT examples of extracting fields from those manuals. It might also need neither step if a prompting baseline already meets the target. Compare these options rather than assuming a mandatory training sequence.

Catastrophic forgetting is degradation of previously learned behavior during adaptation. Mitigations include representative rehearsal data, controlled learning rate and training duration, regularization, and appropriate adapter capacity. No fixed replay percentage or “one tenth of the original learning rate” guarantees retention. Keep the original checkpoint and evaluate old tasks throughout training.

Choose the update and serving strategy

Decision Benefit Cost or limitation
Full-parameter training Can adjust the complete selected representation Gradient and optimizer state, checkpoint storage, and training infrastructure
LoRA / PEFT Smaller trainable state and modular task adapters Frozen base and activations still consume memory; constrained update capacity
QLoRA Quantized frozen base reduces weight storage during adapter training Dequantization, kernel support, and accuracy need checking
Merge an adapter for deployment A simpler fixed serving artifact where supported Loses convenient per-request adapter switching; revalidate quantized merges

A frozen base does not mean the active adapted model preserves every old behavior. Fewer trainable weights also do not mean all computation through the base disappears.

Hyperparameter Tuning

Tune a small, explicit set of experiments before launching a large sweep:

  1. Learning rate and duration: monitor development loss and task outcomes. A high rate may destabilize training; a low rate may under-adapt. Numerical bugs can mimic either.
  2. Batch and sequence length: include long examples from the real workload and account for activation memory.
  3. Adapter rank and target layers: compare capacity and retained-task behavior under a common budget.
  4. Data mixture: inspect whether easy examples dominate important rare cases.
  5. Packing: combine short sequences efficiently only with the intended attention and loss boundaries. An EOS token alone does not prevent cross-example attention.

Pin the tokenizer and chat template used during training and serving. A deployment using the wrong role delimiters can perform poorly even when the training run was sound.

Work the decision with a learning curve

Assume a support task has a fixed development set of 1,000 prompts. These numbers are an interview exercise, not benchmark results.

Candidate Reviewed training examples Correct resolutions Unsupported claims p95 latency
Prompting baseline 0 790 90 2.0 s
Adapter A 500 820 65 2.1 s
Adapter B 2,000 860 30 2.1 s
Adapter C 8,000 862 55 2.1 s

Candidate C is not better merely because it saw more data. Inspect whether new labels introduced unsupported answers or changed the distribution. Candidate B still needs a regression suite, an acceptable refusal rate, and a cost comparison. If preparing its data and maintaining training versions costs more than the measured benefit, keep the simpler baseline.

Interview practice

  1. Is instruction tuning different from SFT? It is a form of SFT focused on following instructions; SFT is the broader supervised adaptation method.
  2. Can fine-tuning replace retrieval? It can encode some information, but freshness, traceable citations, and document-level access control usually require maintained retrieval or other data access.
  3. Why does low training loss not prove success? It measures fit to the selected objective and examples, not independent task performance or retained capabilities.
  4. Why might full training be justified over LoRA? A measured capacity or representation gap may outweigh the extra memory and operational cost. Compare rather than assume.
  5. What is wrong with packing examples using only EOS? EOS is a token, not an attention boundary. The implementation must enforce the intended independent-example behavior.
  6. How do you release the model? Version the full artifact and template, pass task and regression gates, canary with monitoring, and retain a tested rollback.

Recall card and closing

Diagnose → curate → adapt → compare → release. Name the failure that requires changing weights, the evaluation that would disprove your proposal, and the owner of the next data refresh. That is a stronger close than listing training libraries.

Further reading: InstructGPT, QLoRA.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Pretraining: learning a reusable language model
NEXT LESSONLoRA and QLoRA: understand what becomes smaller →

Explore the diagram