Speculative decoding accelerates autoregressive generation by proposing candidate tokens with a cheaper mechanism and verifying them with the target model. An exact sampling algorithm can preserve the target output distribution while accepting several tokens from one verification pass. Speedup is conditional on useful accepted work exceeding drafting and verification overhead.
The Core Concept
Low-batch decode can leave arithmetic capacity underused while moving substantial model data. Verifying several known candidate positions can use that capacity more efficiently than running a separate target step for every new token. At high utilization, extra speculative work may reduce throughput instead. Start from the measured inference bottleneck.
Follow one draft-and-verify round
- The drafter proposes a short sequence conditioned on the accepted prefix.
- The target computes the needed conditional distributions for those candidate positions.
- The acceptance procedure checks candidates in order.
- At the first rejection, it produces a corrected token and discards the dependent later candidates.
- The service commits accepted output, updates the relevant caches, and starts another round.
Read diagram source
flowchart LR
A[Accepted prefix] --> B[Draft candidate block]
B --> C[Target verification]
C --> D{First rejected position?}
D -->|Yes| E[Keep accepted prefix and sample correction]
D -->|No| F[Keep block and use target continuation]
E --> G[Commit tokens and repair cache state]
F --> G
G --> A
For a standard stochastic version, let q(x) be the drafter probability and p(x) the target probability at the same history. A proposed token can be accepted with probability min(1, p(x)/q(x)). On rejection, correction uses a distribution proportional to the positive part of p − q. The exact procedure matters: accepting any “plausible-looking” token does not preserve the distribution. See Fast Inference from Transformers via Speculative Decoding.
For greedy decoding, a compatible procedure instead checks agreement with the target's selected token. Distribution preservation is not a promise of identical random samples across runs, and it does not make the target model factually correct.
Work the latency condition
Assume ordinary target decoding takes 40 ms per committed token. A speculative round takes 15 ms to draft and 45 ms to verify and manage state.
| Average committed tokens per round | Effective time per token | Decision in this exercise |
|---|---|---|
| 4 | 60 / 4 = 15 ms | Faster for this request |
| 2 | 60 / 2 = 30 ms | Smaller benefit |
| 1 | 60 / 1 = 60 ms | Slower than baseline |
These are invented measurements for practice. Real serving also includes queueing, token buffering, drafter memory, cache rollback, and competing requests. Measure accepted tokens per verification, round time, TTFT, inter-token latency, aggregate throughput, and task quality under the actual sampling configuration.
Ways to propose candidates
| Approach | Candidate source | Added cost or constraint |
|---|---|---|
| Separate small model | Its autoregressive predictions | Another model, compatible vocabulary, cache and scheduling |
| Multi-token or extra decoding heads | Predictions at several future offsets | Training and model/runtime-specific support |
| Repeated-prefix or n-gram proposal | Matching token patterns | Useful only when patterns predict the target continuation |
| Lookahead decoding | Parallel candidate generation and verification using the model | More per-step computation and implementation complexity |
Medusa adds heads and uses a tree of candidate continuations. Its training and acceptance variants have different guarantees; “multiple heads” by itself does not prove exact target-distribution preservation. Separate draft models have not become universally obsolete.
Lookahead Decoding uses a parallel decoding formulation related to Jacobi iteration and n-gram candidates. It should not be described merely as searching the model's past hidden states. Pattern-based proposal and model-based lookahead are distinct mechanisms even when both use candidate token sequences.
Adapt speculation to the operating point
A longer draft can amortize verification over more tokens but also costs more to generate and may waste later candidates after an early rejection. At higher batch sizes, the target may already use the accelerator efficiently. Tune or dynamically select draft length only with runtime support and evidence from the workload.
Higher temperature does not mathematically imply low acceptance. Acceptance depends on how closely the drafter and target distributions match under the actual transforms. If p = q, the standard procedure can accept every draft even at a high temperature. Creative tasks may be hard for a given drafter, but temperature alone is not a sufficient diagnosis.
| Failure | Investigation | Repair |
|---|---|---|
| Good acceptance but worse latency | Drafting or verification dominates | Cheaper drafter, shorter block, or disable speculation |
| Benefits disappear at peak load | Extra work competes with useful batching | Use load-aware policy and benchmark throughput |
| Different output behavior | Sampling transforms or acceptance logic differ | Match the target contract and validate exactness where claimed |
| Memory pressure | Second model and candidate caches consume capacity | Account for full state; compare integrated proposal methods |
Interview practice
- What is speculative about the output? Candidates are provisional until target verification accepts them; unverified draft tokens should not be committed as final output.
- Does a rejection waste the whole round? Earlier accepted tokens remain useful. Later dependent candidates are discarded or handled by the particular algorithm.
- Why can exact decoding still produce hallucinations? Exactness preserves the target distribution, including its errors.
- Must the draft model be smaller? It must be sufficiently cheaper for the workload; size is one contributor to cost.
- Does high temperature always defeat it? No. Distribution mismatch and runtime overhead determine the benefit.
- What would make you disable it? Worse useful throughput or latency, unsupported output contracts, poor acceptance, or excessive memory cost.
Recall card and closing
Propose → verify → commit → measure. State the acceptance guarantee, the average useful tokens per round, and the workload range where the optimization pays for itself.