Parameter-efficient fine-tuning (PEFT) adapts a model by training a subset of parameters or additional small components. Low-Rank Adaptation (LoRA) freezes a base weight matrix and learns an additive update represented by two smaller matrices. QLoRA trains adapters through a frozen, quantized base model.
The useful interview distinction is between weight storage, trainable state, activation memory, and serving cost. Reducing one does not automatically reduce all four.
Start with the memory problem
A team needs task-specific versions of one base model for several document workflows. Keeping separate full checkpoints increases storage and deployment work. LoRA can share a base and store smaller task updates. Whether this is worthwhile depends on quality and the serving engine's adapter support.
| Memory component | Ordinary LoRA | QLoRA |
|---|---|---|
| Frozen base weights | Remain present, often at higher precision | Stored in a supported low-bit representation |
| Base gradients and optimizer state | Not needed for frozen weights | Not needed for frozen weights |
| Adapter weights, gradients, optimizer | Required | Required; not necessarily four-bit |
| Activations and temporary buffers | Still required | Still required |
This is why an adapter with very few parameters can still run out of GPU memory on a long training sequence.
LoRA Mechanics
For a column-vector input x, a linear layer becomes:
h = W x + (α/r) B A x
W: d_out × d_in A: r × d_in B: d_out × r
- Compute the ordinary frozen projection
W x. - Project the input into
rdimensions usingA. - Project back to the output dimension using
B. - Scale the update and add it to the frozen projection.
- Update
AandBduring training, leavingWunchanged.
For W = 4096 × 4096 and r = 16, the base has 16,777,216 weights. The adapter has 16 × 4096 + 4096 × 16 = 131,072, or 0.78125% as many. This is a single-layer parameter comparison, not a prediction of whole-model training memory or speed.
Read diagram source
flowchart LR
X[Input x] --> W[Frozen W]
X --> A[Trainable A: reduce dimension]
A --> B[Trainable B: restore dimension]
B --> S[Scale by alpha over rank]
W --> P[Add]
S --> P
P --> Y[Output h]
A common initialization makes one matrix zero so the initial update is zero. α = 2r is one recipe, not a requirement. Rank, scaling, initialization, learning rate, and target layers interact. The original LoRA paper explains the low-rank update formulation.
Select target layers deliberately
Adapting only selected attention projections reduces state. Including attention output and feed-forward projections gives the update more places to change behavior at greater cost. Architecture names differ; inspect the actual modules rather than assuming every model contains q_proj and v_proj.
Run a small comparison with the same data and evaluation: narrow target set versus all supported linear layers, then a rank sweep if needed. For mixture-of-experts models, fused expert parameters and serving support introduce additional compatibility questions. Verify the installed version's PEFT configuration; examples from a development branch may require a newer package than your environment.
QLoRA: separate storage precision from training precision
QLoRA combines a frozen low-bit base with trainable low-rank adapters. Gradients pass through computations involving the base to reach the adapters; the quantized base itself remains frozen.
| Ingredient | Purpose | Limitation to explain |
|---|---|---|
| NF4 | Nonuniform four-bit representation designed around normally distributed weights | Distribution assumptions and block scaling matter |
| Double quantization | Quantizes quantization constants to reduce metadata | Does not remove all scale and metadata overhead |
| Paged optimizers | Manage memory spikes using unified-memory mechanisms | Transfers cost time and do not guarantee freedom from OOM |
The QLoRA paper demonstrated adaptation of a 65B model on one 48 GB GPU with its recipe. Treat that as a documented experiment, not a promise that any 70B workload fits with arbitrary context and batch size. Computation dtype and adapter dtype are separate from the stored base dtype.
Advanced Variants
| Method | What changes | When to investigate |
|---|---|---|
| DoRA | Separates weight magnitude and direction; applies low-rank adaptation to direction | A measured quality gap may justify extra state and computation |
| VeRA | Freezes shared random low-rank matrices and learns scaling vectors | Very small task-specific trainable state is valuable |
| Rank-stabilized LoRA | Uses α/√r scaling rather than α/r |
Testing larger ranks requires stable update scaling |
None is a universal replacement for ordinary LoRA. Compare target quality, retained capabilities, export compatibility, and serving support. Primary descriptions: DoRA, VeRA, rank-stabilized LoRA.
Multi-LoRA Serving
A compatible engine can select an adapter for each request while sharing a base. Define the adapter identity as part of routing and authorization; a customer must not choose another customer's private adapter merely by naming it.
- Pin the base revision, tokenizer, adapter revision, target modules, and supported rank.
- Distinguish adapters stored on disk, cached in CPU memory, and resident on the GPU.
- Measure cold loading and eviction under realistic tenant traffic.
- Batch only combinations supported by the runtime and capacity plan.
- Roll back the adapter and its configuration together.
Merging W + (α/r)BA can produce a fixed dense layer where supported. It simplifies a single-model deployment but sacrifices easy adapter switching. Quantized merges can introduce additional rounding or require conversion; evaluate the actual exported artifact. Consult the current vLLM LoRA serving documentation before relying on specific limits.
Interview practice
- Why can a frozen model still consume training memory? Its weights, forward computations, required activations, and temporary buffers remain; only selected gradient and optimizer state disappear.
- Does rank 16 mean sixteen trainable parameters? No. The two matrices contain
r(d_in + d_out)parameters for one adapted linear layer. - Does QLoRA train four-bit adapters? Not inherently. It quantizes the frozen base; adapter and computation dtypes are separate settings.
- Why not always maximize rank? More capacity costs memory and can overfit. The useful rank depends on data, target layers, and the task.
- Can one base serve a thousand adapters? Storage may be feasible, but active residency, loading, kernel limits, traffic skew, and tail latency determine serving capacity.
- What could break an adapter after a base upgrade? Changed weights, tokenizer/template, dimensions, or module definitions. Compatibility and behavior both need revalidation.
Recall card and closing
Freeze the base. Learn a small update. Measure the whole system. Close with the quality gain, full memory budget, adapter compatibility contract, and latency under cold as well as warm requests. Continue to quantization for the storage and serving tradeoffs.