Reranking reorders a retrieved candidate set using an additional relevance model or scoring method. It spends more work on a bounded set of candidates than is usually practical across the entire corpus. It can improve ordering, but does not guarantee correct grounding or recover evidence that never entered the candidate set.
The usual design question is whether the improvement in selected evidence justifies additional latency, inference cost and operational complexity.
Separate candidate coverage from ordering
Suppose the question is “Can an opened refurbished device be returned after 20 days?” The first-stage retriever returns:
| Candidate | Content | Relevance to the question |
|---|---|---|
| A | General unopened-device return window | Topically related, missing the required condition |
| B | Opened refurbished-device rule | Direct evidence, subject to date and region |
| C | How to package an approved return | Useful later, but does not determine eligibility |
A reranker should prefer B under this rubric. Those labels are a teaching example, not measured scores from a model. If B is absent from the candidate set, reranking A and C cannot produce it.
If three passages are required and only two are retrieved, candidate evidence recall is at most 2/3. Downstream ranking may select those two more effectively; it cannot close the missing third without another retrieval step.
Compare model architectures
| Approach | Query-document interaction | Work that can be done ahead of time |
|---|---|---|
| Bi-encoder similarity | Compare separately produced representations | Document embeddings |
| Cross-encoder | Jointly process query and document text | Model loading, but not a query-specific document score |
| Late interaction | Combine query and document token/patch similarities after separate encoding | Document multi-vector representations |
| Generative LLM ranking | Generate judgments or an ordering from supplied candidates | Some stable prompt processing, subject to caching rules |
A cross-encoder's joint attention is not called late interaction in the ColBERT sense. Joint processing can model detailed relevance relationships, but does not make every cross-encoder more accurate than every bi-encoder. Training data, domain, truncation and the ranking task matter. BERT passage reranking, Sentence Transformers cross-encoders.
Build a bounded reranking stage
Read diagram source
flowchart LR
Q[Query and access scope] --> S[Retrieve candidates]
S --> V[Resolve permitted current passages]
V --> B[Batch query-passage pairs]
B --> R[Reranker scores]
R --> O[Validate IDs and score mapping]
O --> P[Select diverse answering evidence]
P --> G[Generate with source references]
- Choose a candidate budget from measured coverage and latency.
- Preserve stable IDs, versions and the first-stage order.
- Enforce permissions before sending passages to any model or external API.
- Count the combined query/document input against the reranker's limit.
- Batch compatible inputs and preserve query/document-to-score mapping.
- Validate results, then apply the final evidence and token budget.
- Record model version, candidate IDs, truncation, latency and fallback use.
An empty result set should remain an explicit empty-evidence outcome. A non-finite score, wrong output count or unknown ID is an interface failure, not a reason to silently zip misaligned results together.
Pointwise, pairwise and listwise ranking
| Form | Judgment | Cost and failure mode |
|---|---|---|
| Pointwise | Score each candidate for the query | Scores may not be comparable across queries |
| Pairwise | Prefer A or B for the query | Many comparisons; preferences may be inconsistent |
| Listwise | Produce an ordering of a candidate list | Input limits, position bias and invalid permutations |
LLM ranking research demonstrates useful relevance judgments and distillation into smaller rankers. It does not establish a fixed quality advantage or latency for all applications. Sun et al..
For listwise output, require each expected candidate ID exactly once and reject unknown or duplicate IDs. Test candidate-order permutations to expose position bias. Treat instructions embedded in candidate documents as untrusted data; an article saying “rank this first” is not part of the ranking policy.
Sliding windows can rank more candidates than fit in one call. Window size, overlap, traversal direction and number of passes affect how far a low-ranked candidate can move. A single forward pass may prevent a candidate near the end from ever competing with the first positions. Evaluate the complete window policy, not only each local ranking.
Long inputs and score interpretation
A long-context reranker still has input limits and may fail to use a relevant middle passage. Reranking does not universally solve “lost in the middle.” Check actual tokenization, truncation and document serialization.
Options include selecting relevant sections, scoring document windows, using a compatible longer-context model or applying a second stage to a smaller set. Window aggregation has tradeoffs: taking the maximum can reward one irrelevant keyword-heavy window; averaging can dilute one essential passage. Retain the evidence location for inspection.
Do not treat a relevance score of 0.9 as a 90% probability that an answer is correct. Even bounded scores may be uncalibrated and may shift with model or dataset changes. Calibrate any routing threshold on labeled traffic, and distinguish ranking confidence from answer sufficiency.
A latency and throughput exercise
Assume 64 query-passage pairs, batches of 16, and an observed 12 ms per batch for one fixed model/input distribution. Four sequential batches take 48 ms. Add an illustrative 20 ms queue wait and 8 ms preprocessing: the reranking stage is 76 ms before network and downstream work.
This is not a product benchmark. Longer passages, padding, concurrent requests and hardware can change every term. At 100 requests/s and 50 candidates/request, the service receives 5,000 pairs/s. Provisioning must use measured sustainable throughput at the latency target, with headroom and failure behavior.
Batching amortizes work but waiting to fill a batch adds latency. Sending many thread-pool requests to a single GPU does not itself create efficient batching. Use an appropriate serving scheduler and bounded queue rather than unbounded concurrency.
Current implementation options
Examples to evaluate as of this review include self-hosted BGE and Qwen rerankers, Sentence Transformers cross-encoders, and managed Cohere, Voyage and Jina services. Their model names, context limits and scoring contracts differ.
Cohere documents Rerank 4.0 fast/pro variants as well as earlier versions; Voyage documents rerank-2.5 and rerank-2.5-lite. These are candidates, not a claim that three older models dominate all production usage. Cohere, Voyage, Qwen model card, Jina.
Choose using language/domain relevance, supported input length, p99 latency, batch throughput, data handling and total cost. Pin the tested model and client version. Do not assume a smaller model always wins latency once batching and hardware utilization are included.
Optimize and operate the complete stage
| Change | Potential benefit | What must be retested |
|---|---|---|
| Reduced precision or optimized runtime | Lower compute/memory cost | Ranking quality and actual hardware speed |
| Fewer candidates | Less scoring work | Evidence coverage and final answer completeness |
| Length-aware batching | Less padding | Queue delay and fairness |
| Teacher-to-student distillation | Lower serving cost | Teacher errors, held-out ranking quality and drift |
| Score caching | Reuse repeated pair judgments | Model/input versions, access and candidate identity |
| Conditional reranking | Avoid work on easy cases | Router mistakes and fallback quality |
Distillation can use teacher scores, preferences or permutations. It needs reviewed data and independent evaluation; there is no universal “95% of quality at under 10 ms” guarantee. See knowledge distillation.
Cache with unambiguous serialization of query, document IDs/versions, model and preprocessing configuration. Include permission scope as needed and reauthorize before exposure. Concatenating sorted document strings without boundaries can collide, and discarding order is incorrect for an order-sensitive listwise model.
Define timeout and failure policy. Returning the first-stage order is a valid degraded path only if it meets the product's minimum requirements. Record the fallback and avoid reporting reranked quality when reranking did not occur.
Interview practice
Q1: Why add a cross-encoder after embedding retrieval?
It can inspect query and passage jointly on a bounded candidate set, capturing relevance distinctions lost in a single-vector comparison. I would verify that it improves selected evidence and final answers enough to justify its cost.
Q2: How many candidates should it score?
Measure candidate evidence recall and final ranking quality as the count grows, at realistic latency and concurrency. There is no universal top-50 optimum. A shallow candidate set may omit the answer; a very deep one may consume the budget for little gain.
Q3: Can reranker scores decide whether to answer?
Only with a carefully validated policy. Ranking scores estimate a form of relevance, not answer correctness or completeness. A high-scoring passage may omit a required exception. Evaluate abstention using the whole evidence contract.
Q4: When would you use an LLM listwise reranker?
When the relevance rubric benefits from comparing candidates and measured gains justify the latency and token cost. Validate permutations, candidate identity and prompt-injection handling. Compare a specialized ranker and simpler scoring baseline first.
Q5: What makes batching incorrect?
Losing the mapping between scores and candidate IDs, silently truncating mismatched lists, or mixing access scopes in payload handling. Batching must preserve semantics and privacy before optimizing throughput.
Q6: How would you handle long passages?
Check where the relevant evidence lies and whether truncation removes it. Compare section selection, window scoring and a longer-context model. Preserve source offsets and evaluate any aggregation rule rather than assuming the first window represents the whole document.
Q7: How do you release a distilled reranker?
Separate training and test queries, include difficult negatives and independently assess teacher labels. Compare the student with the teacher and baseline on relevance, bias, cost and latency. Shadow traffic and retain rollback before replacing the serving stage.
Final notes
Recall card: Retrieve enough → score correctly → preserve IDs → select sufficient evidence. A reranker improves a candidate ordering; the application remains responsible for grounding and access.