Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Embeddings and Vector Spaces

By Anup Rai50 min readReviewed September 2026

An embedding is a numerical representation of an item in a vector space, usually learned so useful relationships are reflected in the representation. For dense text retrieval, the item is text and the output is typically a fixed-length list of numbers used to compare queries and passages. The learned relationship may be relevance, semantic similarity or another task-specific property.

A vector is an ordered list of numbers. Dense means that typically many of its entries are nonzero. We will unpack both ideas with examples before using them in a search system.

Consider these two sentences:

  • “I forgot my password.”
  • “Steps to reset your login credentials.”

They use different words but express a related need. A useful embedding model produces vectors that score as similar for them. This lets a search system find relevant text even when the wording differs.

This chapter builds that system step by step: what the numbers represent, how a model learns to produce them, how we search with them, and how we check whether the results are useful.

Understand embeddings through a search example

A user asks, “How can I recover access after losing my security key?” The relevant document might say “Account recovery when a hardware authenticator is unavailable.” Exact words overlap only partly. An embedding model maps each text to a numerical representation learned to make useful relationships easier to detect. A retrieval system compares the query representation with compatible document representations and returns candidates.

The important word is learned. An embedding does not contain a universal measure of meaning. Training examples and objectives decide which relationships it tends to preserve. A model trained to match search queries with answers can behave differently from one trained to group similar sentences. Two texts can be close because they discuss the same topic even when one denies the other's claim. Similarity therefore helps find evidence; it does not establish that the evidence is true, current, authorized, or sufficient.

Follow the pipeline before tuning dimensions

Start by choosing the unit to retrieve: a paragraph, a procedure, a table row with its headers, or another coherent evidence unit. Apply the model's documented query/document formatting and preprocessing. Store vectors with text, document version, and permissions. At query time, enforce access scope, search for candidates, and optionally use a richer reranker. Measure whether the needed evidence is actually retrieved before judging the generated answer.

A bi-encoder computes query and document representations separately, so document vectors can be prepared ahead of time. A cross-encoder examines a query and candidate together, allowing richer interaction at greater per-pair cost. ColBERT-style late interaction keeps multiple token representations and compares them at search time. These are different ways of trading stored representation, query-time work, and relevance; no fixed latency number decides the right choice for every corpus.

Two kinds of recall that are easy to confuse

ANN recall asks whether approximate nearest-neighbor search recovers the neighbors that exact vector search would have found. Task retrieval recall asks whether it finds the documents or passages needed to answer the real question. You can have perfect ANN recall and poor task recall if the embeddings rank the wrong content highly. Tune index parameters only after checking whether the representation and corpus contain the evidence you need.

For memory, keep this chain: choose the evidence unit → create compatible representations → search within permissions → judge relevance → use evidence. The mathematical sections below explain similarity, dimensions, training losses, and indexes; this chain explains why those details matter.

Table of Contents

What an Embedding Represents

An embedding model accepts text and returns numbers:

"Reset my password" → embedding model → [0.12, -0.34, 0.81, ...]

These numbers are illustrative. We do not assign them by hand; they are the model's output.

Dimensions: how many numbers are in the list?

[0.12, -0.34, 0.81] has three numbers, so it is a three-dimensional vector. A 1,024-dimensional embedding has 1,024 numbers.

For a given model and output setting, a short question and a long paragraph both produce the same number of coordinates. A longer input does not automatically produce a longer vector. It must still fit the model's input limit.

Three quantities are easy to confuse:

Quantity What it counts Example
Input tokens Pieces of text the model reads; a piece can be a word, part of a word, or punctuation A passage might contain 500 tokens
Model parameters, also called weights Internal numbers learned during training A model might have 600 million parameters
Embedding dimensions Numbers in the output vector for one item That model might output 1,024 numbers per passage

What is a vector space?

With two numbers, [3, 4], we can plot a point three units across and four units up. We can also draw an arrow from the origin [0, 0] to that point.

A vector space is a set of vectors with addition and scalar multiplication satisfying the vector-space rules, including associativity, distributivity, a zero vector and additive inverses. Adding its vectors or multiplying them by allowed scalars stays in the space. For ordinary real-valued embeddings, the surrounding space is Rd\mathbb{R}^d: all lists of d real numbers.

For example, [1, 2] + [3, 4] = [4, 6], and 2 × [1, 2] = [2, 4]; both results remain in R2\mathbb{R}^2. To compare lengths, angles or distances we additionally choose a norm, inner product or distance rule. The familiar dot product and Euclidean length are common choices for embeddings, not part of every abstract vector space's definition.

The set of embeddings a particular model actually produces need not fill this space or itself be closed under addition. Adding two vectors is valid arithmetic without guaranteeing that the result describes a meaningful sentence.

The coordinates usually do not have human-readable labels. Coordinate 17 does not reliably mean “finance.” Meaning is represented by patterns across many coordinates. Models learn which patterns help with their training task.

Dense versus sparse: a dense embedding might have 1,024 mostly nonzero numbers. A sparse representation might have a slot for every term in a large vocabulary, with nonzero weights in only a few slots. Both are vectors; they organize information differently. We will return to sparse representations when combining word matching with semantic search.

What the numbers can and cannot tell us

An embedding comparison estimates a learned relationship. For search, we want that relationship to be relevance: “Does this passage help answer this question?” Two useful texts need not be paraphrases.

Closeness does not establish truth. “Payment succeeded” and “payment failed” discuss the same topic but contradict one another. A model can put them close together, so applications must check whether its distinctions are good enough.

Embeddings can also support clustering (grouping related items), deduplication (finding repeated or near-repeated items), classification (assigning categories), and recommendations. A model that works well for search may not work equally well for each of these tasks.

The historical analogy king − man + woman ≈ queen illustrates a pattern found in some word-vector spaces. It is not a general reasoning rule. Likewise, squeezing hundreds of dimensions into a two-dimensional plot can distort distances; attractive clusters do not prove search quality.

You may see the definition written as:

fθ(x)∈Rd f_\theta(x)\in\mathbb{R}^d

Read it as: “The model, f, with learned weights θ, turns input x into a list of d real-valued numbers.” It is mathematical shorthand for the text-to-vector operation above.

One Example from Query to Answer

Throughout the chapter, we will use this search question, or query:

“I lost my phone. How can I sign in?”

Our collection of documents—called the corpus—contains three passages:

Passage Source text Does it address the question?
A “Use a backup recovery code when your authenticator device is unavailable.” Yes: it provides another way to sign in
B “Report lost company phones to IT.” Related topic, but it does not explain how to sign in
C “Change your profile photo in account settings.” No

A keyword search might be attracted to B because it contains “lost” and “phones.” We want semantic search to recognize why A is useful despite its different wording.

Before anyone searches

  1. Break long documents into manageable passages. These pieces are called chunks.
  2. Run each passage through the embedding model. This operation is called encoding.
  3. Store each vector with its passage ID. Keep the original text, source location, and access permissions too.

We do this work ahead of time because the same passage may be useful for thousands of future queries.

When a query arrives

  1. Encode the query into a vector that can be compared with the stored passage vectors.
  2. Calculate similarity scores and retrieve the best candidates the user is allowed to access. “Retrieve” simply means find and return.
  3. Fetch their source text. Optionally use a second model to check the candidates more carefully and reorder them; this is reranking.
  4. Return the passages, or give the selected text to a language model so it can answer using that evidence.

That final combination is retrieval-augmented generation (RAG): retrieve evidence first, then generate an answer using it. The language model normally receives the retrieved text, not the embedding numbers.

Architecture / visual model
flowchart TD D[Prepare passage vectors and store source text] --> S[Search passages the user may access] Q[User asks a question] --> E[Convert question into a vector] E --> S S --> T[Fetch candidate passage text] T --> R[Optionally rerank the candidates] R --> A[Return evidence or use it to generate an answer]
Read diagram source
flowchart TD
    D[Prepare passage vectors and store source text] --> S[Search passages the user may access]
    Q[User asks a question] --> E[Convert question into a vector]
    E --> S
    S --> T[Fetch candidate passage text]
    T --> R[Optionally rerank the candidates]
    R --> A[Return evidence or use it to generate an answer]

A search result is only a candidate. Even when no passage answers a question, one passage can still score highest. Later, we will evaluate when the system should say “I could not find a supported answer.”

Comparing Vectors

To rank passages, we need a rule that converts two vectors into a score. Three common rules are dot product, cosine similarity, and Euclidean distance.

Use these invented vectors to make the arithmetic small enough to follow. They are not outputs from a real model:

Query q = [1, 0]
Passage A = [0.8, 0.6]
Passage B = [0.6, 0.8]
Passage C = [0, 1]

Dot product: multiply matching entries, then add

For the query and A:

q · A = (1 × 0.8) + (0 × 0.6) = 0.8

Doing the same for B and C gives 0.6 and 0. The highest score wins, so the ranking is A, B, C.

For longer vectors, we repeat that operation across every coordinate:

a⋅b=∑iaibi a\cdot b=\sum_i a_i b_i

The symbol Σ means “add these terms”; i identifies a coordinate. No new operation is hiding in the formula.

Length and normalization: separate size from direction

The vector [3, 4] has length 5, from Pythagoras: √(3² + 4²) = 5.

Dividing every coordinate by 5 gives [0.6, 0.8]. This arrow points in the same direction, but its length is now 1. This operation is L2 normalization; the result is a unit vector.

Why care? Replace B's vector [0.6, 0.8] with [3, 4]. Its dot product with the query jumps from 0.6 to 3, even though its direction did not change. A dot product is influenced by both direction and vector length, also called magnitude.

Cosine similarity: compare directions

Cosine removes the effect of length by dividing the dot product by both vector lengths:

cos⁡(a,b)=a⋅b∥a∥∥b∥ \operatorname{cos}(a,b)=\frac{a\cdot b}{\lVert a\rVert\lVert b\rVert}

The notation ‖a‖ means “the length of a.” For our query and [3, 4], cosine is 3 / (1 × 5) = 0.6. Scaling B no longer gives it an advantage.

  • 1: the vectors point in the same direction.
  • 0: the vectors meet at a 90-degree angle (perpendicular); the score alone does not prove the texts are unrelated.
  • −1: the vectors point in opposite directions.

Cosine ranges from −1 to 1 for nonzero real vectors, including normalized vectors. A zero vector has no direction, so its cosine is undefined and must be handled explicitly.

Euclidean distance: measure the gap between points

Euclidean distance is straight-line distance. Subtract matching coordinates, square the differences, add them, then take the square root:

∥a−b∥2=∑i(ai−bi)2 \lVert a-b\rVert_2=\sqrt{\sum_i(a_i-b_i)^2}

Using the original example vectors, the distance from the query to A is √((1 − 0.8)² + (0 − 0.6)²) = √0.4 ≈ 0.632. Here smaller is better. Distances to B and C are approximately 0.894 and 1.414, so A still wins.

Which rule should I use?

Rule What to remember Best match
Dot product Multiply and add; length influences the result Highest score
Cosine Compare direction while ignoring length Highest score
Euclidean distance Measure separation Lowest distance

If both vectors have length 1, cosine equals dot product. Their squared Euclidean distance is:

∥a−b∥22=2−2(a⋅b)=2−2cos⁡(a,b) \lVert a-b\rVert_2^2=2-2(a\cdot b)=2-2\operatorname{cos}(a,b)

As cosine increases, distance decreases. Therefore, all three give the same exact ranking for unit vectors. Search systems that use approximations can still return different candidates.

Follow the embedding model's documented metric and normalization recipe. If it learned to use magnitude, removing magnitude changes the scores. Also check database conventions: a field named “distance” may contain 1 − cosine or squared Euclidean distance.

A similarity of 0.8 does not mean an 80% chance of relevance. The score needs evaluation on real queries before it can support a decision such as accepting or rejecting a result.

How Embedding Models Learn

We have seen how to compare numbers. The next question is: how does a model learn to produce numbers that make useful passages score highly?

It learns from examples of what should match. For our lost-phone query, a training example could identify the backup-code passage as useful and the profile-photo passage as unhelpful.

Follow one training step

  1. Create vectors. The model encodes the query and both passages using its current weights.
  2. Compare them. Calculate a similarity score for each query–passage pair.
  3. Measure the mistake. A mathematical rule called a loss function gives a penalty when the useful passage scores too poorly compared with the unhelpful one.
  4. Adjust the model. Backpropagation calculates how a small change to each weight would change that penalty. An optimizer uses those calculations to make small weight updates intended to reduce it.

Repeat across many examples. The model can learn patterns that help it match new questions with useful passages, including wording it has not seen together before. A single update does not guarantee that every example improves.

We do not manually edit every output vector. We update the model that produces the vectors. Most applications start with an already-trained model; training it further on their own examples is called fine-tuning.

Ordinary search uses the trained weights without changing them. This use of a trained model is called inference. Encoding new documents is also inference; it is not automatically another training step.

Positives and negatives: names for the examples

This approach is called contrastive learning because training contrasts useful matches with unhelpful matches.

Term Plain meaning Example for the lost-phone query
Positive A useful match The backup-code instructions
Easy negative An obviously unhelpful match A cake recipe
Hard negative A plausible-looking but unhelpful match Reporting a lost company phone to IT
False negative A useful passage mistakenly treated as unhelpful Another valid way to recover account access

Hard negatives teach the model that shared words or topics are insufficient: the passage must address the question. We can find them among high-ranking but irrelevant search results. False negatives teach the wrong lesson, so training labels need checking.

A batch is a group of examples processed together. Suppose a batch contains a password-recovery question and a cooking question, each with its answer. The cooking answer can serve as a negative for the recovery question. These in-batch negatives reuse work already being done, instead of encoding a separate set of negatives for every question. Check for duplicates and overlapping answers; another example's answer is not always irrelevant.

Where do the training examples come from?

Source What it can teach What needs checking
Human-rated query–passage pairs Which passages answer a query Judgment consistency and domain coverage
Search logs Which results people selected People click prominent results, not only useful ones
Titles paired with article bodies Topic relationships A title may not describe every paragraph
Paraphrases and translations Different ways to express similar meaning Meaning may differ in small but important ways
Model-generated questions and answers Additional examples at scale Generated labels and answers can be wrong

Natural-language inference (NLI) datasets label whether one statement follows from another, contradicts it, or is undecided. Such labels can help train representations. “A dog runs” supports “An animal runs,” but the two are not exact paraphrases; the training relationship matters.

Some training examples have several correct answers. Supporting multiple positives and removing duplicate examples can reduce the chance of treating a correct answer as a negative.

What the contrastive loss formula means

One common loss first converts the candidate scores into a probability distribution:

P(positive∣q)=exp⁡(s(q,d+)/τ)exp⁡(s(q,d+)/τ)+∑jexp⁡(s(q,dj−)/τ) P(\text{positive}\mid q)= \frac{\exp(s(q,d^+)/\tau)} {\exp(s(q,d^+)/\tau)+\sum_j\exp(s(q,d_j^-)/\tau)}

Read the symbols in this order:

  • q is the query; d⁺ is the positive passage; dⱼ⁻ are the negative passages.
  • s(q, d) is their similarity score.
  • τ, pronounced “tau,” is a positive number called training temperature.
  • exp turns each scaled score into a positive number, with larger scores producing larger values. The denominator adds those numbers for all candidates. Dividing by that total gives each candidate a share between 0 and 1.

This conversion is called softmax. If the resulting shares are 0.7 for the positive and 0.2 and 0.1 for two negatives, the positive receives most of the share. These are probabilities within this training candidate set. They do not directly tell us how likely a passage is to be relevant for a real user.

The loss is:

L=−log⁡P(positive∣q) \mathcal{L}=-\log P(\text{positive}\mid q)

Here log is the natural logarithm. A positive share of 0.7 gives a loss of about 0.357; a share of 0.1 gives about 2.303. Lower is better, so training rewards giving the positive more of the share. Code normally uses a stable log_softmax or logsumexp operation to avoid overflow from large exponentials.

A lower temperature makes score differences more pronounced in softmax; a higher one spreads the shares more evenly. It affects the training signal. It does not change how creative a generated answer is—that is a separate generation-time use of temperature.

Distillation is another training technique: a stronger teacher model supplies relevance scores or rankings, and a smaller model learns to imitate them. The purpose is to retain useful judgments while making serving cheaper.

Why query and document instructions matter

A question and its answer play different roles. “How can I sign in?” is useful alongside “Use a recovery code,” although their meanings are not identical.

Some models are trained to follow task descriptions such as “represent this question for finding an answer.” This is instruction tuning for embeddings: the instruction helps specify which relationship the vector should capture.

A related input requirement is a query/document role label. For example, E5-large-v2 expects query: before a query and passage: before a document. A role prefix is not the same as support for arbitrary task instructions. Follow the model's own format instead of inventing a universal prefix. E5 model card.

The query and document can use the same underlying model with different prompts, or separate encoders trained to work together. Cosine compares two vectors symmetrically, but producing those vectors can depend on their roles. Swapping the raw texts between query and document roles can therefore change the final retrieval score.

Remember the learning loop: examples → vectors → scores → loss → weight updates.

From Words to a Passage Vector

Training explains how a model improves. Now look inside the operation that turns text into a vector.

Earlier models gave each word a fixed representation

In a static word model, “bank” has the same vector in “river bank” and “bank account.” The surrounding sentence does not change it.

Three important historical approaches are:

  • Word2Vec: learns from nearby words. Its skip-gram method predicts surrounding words from a selected word; CBOW, or continuous bag of words, predicts a word from its surroundings.
  • GloVe: learns from counts of how often words occur together across a collection of text.
  • FastText: includes character fragments within words. This helps it construct representations for rare words or words absent from its training vocabulary. The standard word vectors remain static.

These methods established that useful relationships can be learned in vector form. Their fixed treatment of a word is a limitation when its meaning depends on context.

Modern models let surrounding text change the representation

A tokenizer splits the input into tokens. A Transformer then lets token representations incorporate information from surrounding tokens, subject to the model's attention rules. You do not need the attention equations here: the important result is that “bank” can acquire different vectors in the two sentences above.

These are contextual token vectors: one representation for each token, influenced by the text it appears in.

Search often needs one vector for the whole passage. We therefore need another step to combine or select information from the token vectors.

Pooling: turn several token vectors into one passage vector

Pooling is that combination or selection step. A tiny mean-pooling example shows the mechanics:

Token vector 1: [1, 2]
Token vector 2: [3, 4]
Average:       [(1 + 3)/2, (2 + 4)/2] = [2, 3]

These are invented numbers. A real model uses contextual token vectors and its trained pooling method.

Method How it gets one vector Important detail
Mean pooling Average the selected token vectors coordinate by coordinate Exclude padding tokens added only to make batch lengths equal
CLS pooling Use the vector of a designated special token, often written [CLS] That token must have learned to carry information useful for the task
Last-token pooling Use the final eligible token's vector Common in decoder-based models, where later tokens can incorporate earlier context
Learned pooling Train a mechanism to choose or weight information from the tokens It must be trained along with the representation it produces

A model may also apply a projection, a learned transformation that changes the vector's coordinates or dimension, and then normalize the result.

Text → tokens → contextual token vectors → pooling → optional projection/normalization

Compressing a passage into one vector can lose detail. Training helps preserve the distinctions needed for the task. Taking an arbitrary language model's hidden vectors and averaging them is therefore not guaranteed to produce a good search model. Use the embedding model's documented recipe. Sentence-BERT paper.

Retrieval Architectures

There are several ways to compare a query with a passage. The main design question is when the query and passage interact.

Bi-encoder: represent each text separately

In a bi-encoder, we encode the query and passage independently, then compare their vectors:

Passage → encoder → passage vector ┐
                                  ├→ similarity score
Query   → encoder → query vector   ┘

“Bi” refers to the two encoding paths; they may share the same model weights. The common single-vector version stores one vector per passage.

Because a passage vector does not depend on the current query, we can compute and store it in advance. Each incoming query then needs only its own encoding plus vector search. This is what makes the approach practical for large collections.

Cross-encoder: read the query and passage together

A cross-encoder receives the pair as one input and produces a relevance score:

[query text + passage text] → model → relevance score

It can examine details of how this particular passage answers this particular question. That can improve ranking, but it needs new model work for each query–passage pair. It cannot reuse a single precomputed passage vector as the complete pair score.

For a million passages, reading every pair is usually too expensive. A practical design is to retrieve, say, 100 candidates with a bi-encoder, fetch their text, then use a cross-encoder to select the best 10. These counts are tuning choices, not rules. Cross-encoder documentation.

Reranking improves the order of candidates already found. It cannot recover an answer missing from that candidate set.

ColBERT: keep token detail, but encode separately

One passage vector must summarize everything in the passage. ColBERT retains multiple contextual token vectors instead. It still encodes queries and documents independently, allowing document representations to be stored ahead of time.

At search time, it compares the query's token vectors with the passage's token vectors. This is called late interaction: the two sides meet after encoding, during scoring.

The original ColBERT scoring rule works as follows:

  1. For each query token, compare its vector with all eligible document-token vectors.
  2. Keep the highest similarity for that query token: its maximum similarity, or MaxSim.
  3. Add those best scores across query tokens.

If two query tokens have best matches of 0.9 and 0.8, the total is 1.7. It is a ranking score, not a probability. The same document token may be the best match for several query tokens.

s(q,d)=∑imax⁡j(qi⋅dj) s(q,d)=\sum_i\max_j(q_i\cdot d_j)

Here i goes through query tokens and j through document tokens. Original ColBERT projects and L2-normalizes token vectors, so each dot product compares directions. ColBERT paper.

Keeping token vectors usually costs more storage and scoring work than keeping one vector per passage. Test whether the extra detail improves your queries enough to justify that cost. ColBERT-style systems can retrieve with suitable indexes (structures that organize the stored vectors for search) or rescore a shortlist; there is no universal storage multiplier or latency requirement.

Method What is stored for a passage? Work when a query arrives
Single-vector bi-encoder One reusable vector Encode query, then compare vectors
Cross-encoder Candidate text is needed Read each query–passage pair together
ColBERT-style late interaction Multiple reusable token vectors Encode query, then compare token vectors

Why combine embeddings with word matching?

Suppose the query contains the error code AUTH-104. An exact identifier may matter more than general semantic similarity. Word-based retrieval can complement dense embeddings here.

BM25 ranks documents using matching terms. It gives more weight to informative terms that occur in fewer documents, limits the benefit of repeating a term, and accounts for document length. It does not need dense embeddings.

Learned sparse retrieval also uses mostly zero term-weight vectors, but a model chooses the weights. It may assign a weight to a related term that was absent from the original text, helping bridge some wording differences.

Hybrid retrieval runs lexical (word-based) and dense search, combines their candidate lists, removes duplicates, and optionally reranks. For identifiers, configure how the search engine splits text—its analyzer—or use a field that preserves the entire identifier. BM25 cannot enforce an exact match on a code that preprocessing has broken apart. Tokenizer reference.

Combining lists with reciprocal rank fusion

BM25 and cosine scores use different scales, so directly adding them may let one dominate simply because its numbers are larger.

Reciprocal rank fusion (RRF) combines positions in the ranked lists instead. A high position contributes more than a low one:

RRF⁡(d)=∑r1c+rank⁡r(d) \operatorname{RRF}(d)=\sum_r\frac{1}{c+\operatorname{rank}_r(d)}

For a document d, look at its rank in each result list r. Ranks start at 1. The constant c reduces how strongly the top few positions dominate; a document absent from a list contributes zero for that list.

For example, with c = 60, a passage ranked first in lexical search and third in dense search gets 1/61 + 1/63 ≈ 0.0323. Calculate this for each candidate, then sort by the combined score. Evaluate the result on your workload. RRF paper.

Chunking and Context

A document may contain many topics, but a search query usually needs a particular part. Chunking splits the document into the pieces we embed and retrieve.

Why chunk size matters

Imagine a handbook covering sign-in, billing, and profile settings. One vector for the entire handbook may blur those topics. Smaller passages give search more focused targets and make it easier to cite the relevant source.

But splitting too aggressively loses meaning:

Recovery codes let you sign in without your phone. Each code can be used once. They stop working after revocation.

The final sentence alone leaves “They” unexplained. A useful chunk should retain enough context to identify the subject.

Start with coherent paragraphs or sections. Keep headings with their content, table headers with their rows, and code blocks intact where possible. Sizes such as 300–600 tokens are experiments to evaluate, not universal recommendations.

The model's context window is the maximum input it can process at once. Count tokens with its tokenizer, including instructions and special tokens. Detect truncation, where text beyond a limit is dropped; silently losing the end can remove the answer.

Overlap and parent text

Overlap repeats some text in adjacent chunks. This can preserve a sentence that crosses a boundary, at the cost of extra embedding work, storage, and duplicate search results.

Store a chunk's ID, its parent document ID, source location, version, and permissions. After finding a small relevant chunk, you can fetch nearby text or the containing section when the answer needs more context.

A document short enough to fit the model does not have to be split. Still, smaller passages may improve retrieval focus; fitting the input limit and being a good search unit are different questions.

Late chunking: read the context before making passage vectors

Ordinary chunking usually follows this order:

Split document → encode each chunk separately → pool each chunk's token vectors

Once the final sentence above is separated, its encoder cannot see the earlier explanation of “They.”

Late chunking changes the order:

Encode a larger span together → choose chunk boundaries → pool tokens within each chunk

The model first produces contextual token vectors while the surrounding text is available. Pooling then creates a vector for each chosen chunk. The “They” tokens can therefore carry information about recovery codes.

This requires access to contextual token vectors and suitable pooling, or a service interface that explicitly supports late chunking. It does not overcome the model's context limit, and a token can use only context allowed by the attention pattern. The original work appeared in 2024; there is no intrinsic 8,000-token minimum. Late Chunking paper.

The similar names describe different operations:

Technique What happens later? Typical result
Late chunking Pooling into chunk vectors happens after a larger span is encoded One vector per chunk can still be enough
Late interaction Query and document token vectors meet during scoring Multiple token vectors are retained

Another approach, contextual enrichment, prepends a title or short explanation to a chunk before embedding it—for example, “Account access: recovery codes.” Check any model-generated explanation for accuracy and preserve the original source text separately.

Generate Real Embeddings

We can now connect the concepts to code. This example uses a trained Qwen model through the Sentence Transformers Python library. Use a tested compatible combination of current PyTorch, sentence-transformers and Transformers releases. The model card documents Qwen3 architecture support from Transformers 4.51.0; that minimum is not a recommendation to select an old release. The first run downloads the weights. It performs inference, not training. Qwen model card.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B")
query = "I lost my phone. How can I sign in?"
passages = [
    "Use a backup recovery code when your authenticator device is unavailable.",
    "Report lost company phones to IT.",
    "Change your profile photo in account settings.",
]

# Prepare one unit-length vector per passage.
doc_vectors = model.encode(passages, normalize_embeddings=True)

# This model supplies a query prompt; passages do not use that prompt.
query_vector = model.encode(
    [query], prompt_name="query", normalize_embeddings=True
)[0]

# Compare every passage with the query, then print highest scores first.
scores = doc_vectors @ query_vector
for index in scores.argsort()[::-1]:
    print(f"{scores[index]:.3f}: {passages[index]}")

Read the important lines as follows:

  1. model.encode(passages, ...) returns a row of numbers for each passage. In an application, store these rows instead of recomputing them for every query.
  2. normalize_embeddings=True makes each vector length 1, so dot product equals cosine.
  3. [query] sends a list containing one query; [0] takes its single returned vector.
  4. @ calculates the dot product between each passage row and the query, producing three scores.
  5. argsort() returns positions ordered from lowest score to highest. [::-1] reverses that order. We use each position to print its score and original passage.

The numbers depend on the actual model run; the earlier two-dimensional vectors were teaching examples. This code compares every passage, which is simple for three items. Larger collections need the search methods below. A deployed application also needs access checks, error handling, and fixed, tested library and model versions. Sentence Transformers API.

Storage and Search at Scale

Three passages are easy to score. Millions of passages create two costs: storing their vectors and finding good matches quickly.

Calculate the storage before choosing an optimization

Suppose we have 10 million chunks, each with a 1,024-dimensional vector. If each number uses four bytes, one vector needs 1,024 × 4 = 4,096 bytes.

Across all chunks:

raw vector bytes=chunk count×dimensions×bytes per number \text{raw vector bytes}=\text{chunk count}\times\text{dimensions}\times\text{bytes per number}

That gives 10,000,000 × 1,024 × 4 = 40,960,000,000 bytes, or 40.96 GB using decimal units. This is just the vectors. Source text, metadata, search structures, backup copies, and replicas—copies serving availability or traffic—need additional storage.

There are two distinct ways to shrink vectors: keep fewer numbers, or use fewer bits to store each number.

Matryoshka embeddings: keep a shorter useful vector

If a model returns 1,024 numbers, can we just keep the first 256? Only if the model supports that use. Arbitrary truncation can discard important information.

Matryoshka Representation Learning (MRL) trains several vector lengths together. The name refers to nested dolls: smaller representations fit inside the larger one. During training, the model is encouraged to make the first 256 numbers useful, the first 512 useful, and the full vector useful, for example.

With a supporting model, reducing 1,024 dimensions to 256 uses one quarter of the raw vector storage and fewer arithmetic operations per comparison. Retrieval quality may change, so measure it rather than assuming a fixed percentage loss. MRL paper.

Use the model's output-dimension option or its documented slicing procedure. For unit-vector scoring, a manually shortened vector needs normalization again: the shorter list generally no longer has length 1. Some models require additional preprocessing, so follow their recipe rather than treating slicing as a universal procedure.

A two-stage search might find candidates using 256 dimensions, then rescore that shortlist using full vectors. It saves work on the broad search, but storing full vectors still costs space, and rescoring cannot recover a passage that the first stage missed. Smaller output vectors also do not necessarily make the encoder faster or reduce token-based API charges.

Quantization: store each value with less precision

Instead of reducing the number of coordinates, we can store a rougher version of each coordinate. This is quantization. Think of keeping fewer possible numeric levels, with a scale to interpret them.

Format How numbers are stored Raw total for our example
FP32 32-bit floating-point numbers: four bytes each 40.96 GB
FP16 16-bit floating-point numbers: two bytes each 20.48 GB
INT8 scalar quantization Map each coordinate to one of 256 integer codes: one byte each 10.24 GB
Packed binary quantization One bit per coordinate, packed eight bits per byte 1.28 GB

These totals omit scales, lookup tables, indexes, and any retained full-precision vectors. Compression can change rankings; its quality loss depends on the model, data, and search method.

Scalar quantization compresses individual coordinates. Binary quantization can reduce each to a bit, for example by recording its sign. Binary codes can be compared using Hamming distance, the number of differing bits, but that distance does not generally preserve the original cosine ranking.

Product quantization (PQ) compresses groups of coordinates. It divides a vector into smaller pieces and represents each piece by the ID of a learned representative pattern. Those patterns form a codebook, a lookup table. Scoring then uses the representatives as approximations to the original pieces.

A common design searches compressed vectors first and checks the candidates using more precise vectors. Evaluate both missed candidates and final ranking quality.

This section compresses output vectors. Quantizing the model's weights is a separate choice that reduces the memory needed to run the encoder and can affect its speed and outputs.

Exact search: find the true best matches

The code above scores every passage and sorts the results. This exhaustive approach guarantees the best matches under its chosen score and stored representation.

Exact search means guaranteeing those nearest neighbors. Exhaustive comparison is one way to do it; some exact algorithms can safely rule out candidates without scoring all of them.

Comparing two d-dimensional vectors needs work proportional to d. Comparing a query with N vectors therefore needs roughly N × d work, written O(Nd). Here N is the number of vectors and d their dimension; the notation describes how work grows, not a measured response time.

Approximate search: trade some accuracy for efficiency

With millions of vectors, we may accept occasionally missing a true nearest neighbor to reduce time or memory. This is approximate nearest-neighbor search (ANN).

An index is a data structure that organizes vectors for search. ANN indexes may skip unlikely candidates or use compressed scores. Even scanning every PQ code can be approximate relative to the original uncompressed vectors. Faiss index reference.

Two common ways to organize candidates are:

  • HNSW (Hierarchical Navigable Small World): connect nearby vectors in a graph, with upper layers providing longer-range navigation. Search follows promising connections toward nearby candidates. The efSearch setting controls how broadly it explores: increasing it generally finds more of the true neighbors but takes more work. The graph itself needs memory.
  • IVF (inverted file): group vectors around representative centers. A query searches selected groups instead of the whole collection. nprobe controls how many groups are searched. Searching more groups generally reduces misses while increasing work.

IVF-PQ combines these ideas: search selected IVF groups and score their vectors using PQ compression.

No index has a universal constant or logarithmic response time for every workload. Measure your corpus, dimensions, hardware, simultaneous queries, and permission filters. The question is: how much speed or memory do we gain, and how many useful candidates do we lose?

Choosing a Model

A model specification tells you what can run. Evaluation tells you whether it solves your search problem.

Start with a small set of realistic questions and known useful passages. Compare candidate models using the same documents and evaluation rules. The next section explains the metrics.

Read the specifications correctly

  • Output dimensions: numbers per vector. More dimensions mean more storage; they do not automatically mean better retrieval.
  • Input capacity: how much text the model can read in one input. A large limit does not guarantee good retrieval from one giant chunk.
  • Hosted API (application programming interface): your program sends requests to a provider that runs the model. Consider price, reliability, data handling, and input/batch limits.
  • Open weights: you can obtain the trained parameters and run the model yourself, subject to its license. You take responsibility for hardware, serving, and updates.
  • Multilingual: supports multiple languages. Test your languages, including whether a query in one language can retrieve an answer in another.
  • Multimodal: can represent more than text, such as images or audio. Check the supported input combinations; some workflows instead extract text from images using OCR, or optical character recognition.

Comparing vectors requires compatible representations. Two models can both output 1,024 numbers but assign very different meanings to their coordinates. Matching vector length alone does not make their outputs comparable. Cross-language and cross-modal retrieval also require a model trained to align those representations.

Use benchmarks to make a shortlist

The Massive Text Embedding Benchmark (MTEB) evaluates models across tasks such as retrieval, classification, and clustering. Its overall score summarizes several abilities; it is not automatically a score for your retrieval problem.

Compare the same benchmark version, tasks, and metrics. Then test your domain: account-recovery questions, exact error codes, confusing alternatives, and questions without answers. A leaderboard cannot choose chunk sizes, acceptable latency, or data-handling requirements for you.

Model reference: snapshot verified 24 September 2026

Use the table to look up candidates after deciding your requirements. It is not a ranking. Input capacity and output dimension are different quantities; “k” denotes thousands of tokens, with the exact limit defined by the provider. Providers may also impose request, batch, and modality-specific limits.

Model or family Output dimensions Documented input capacity Useful distinction
OpenAI text-embedding-3-small / text-embedding-3-large 1,536 / 3,072 default; configurable shortening 8,192 tokens Hosted text embeddings; established baseline. Docs
Voyage voyage-4-large, voyage-4, voyage-4-lite 1,024 default; 256, 512, 2,048 options 32,000 tokens Text retrieval family with documented shared-space compatibility. Docs
Voyage voyage-code-4 / voyage-4-nano 1,024 default; 256, 512, 2,048 options 32,000 tokens Code-focused hosted model / open-weight member of the compatible general-purpose Voyage 4 space. Do not infer code-model compatibility from the shared number alone. Docs
Cohere embed-v4.0 1,536 default; 256, 512, 1,024 options 128k tokens Text, images, and mixed document content. Docs
Google gemini-embedding-2 3,072 default; flexible 128–3,072, with 768/1,536/3,072 recommended 8,192 tokens; additional modality limits Stable multimodal model for text, image, audio, video, and PDF inputs. Model page
Qwen3-Embedding 0.6B / 4B / 8B Up to 1,024 / 2,560 / 4,096 respectively 32K tokens Open-weight text models across compute budgets; instruction-aware and MRL-capable. Model card
Qwen3-VL-Embedding 2B / 8B Up to 2,048 / 4,096 respectively 32K tokens Open-weight text/image/video retrieval, including mixed inputs. Model card
BGE-M3 1,024 dense dimensions 8,192 tokens Established multilingual model supporting dense, sparse, and multi-vector retrieval. Model card

Older baselines include BGE-large-en-v1.5, E5-large-v2, GTE-large, and Nomic-embed-text-v1.5. They remain useful comparison points; Nomic v1.5 also supports Matryoshka representations. Select using current model cards and your evaluation, rather than age alone.

Voyage documents a shared space within the Voyage 4 family. Compatible settings can allow a larger document encoder and cheaper query encoder; this is a documented exception, not a property of arbitrary models. Voyage 4 announcement.

Evaluation and Operations

A search service that returns results is not enough. We need evidence that it finds useful passages, stays fast enough, and keeps results correct as documents and models change.

Build an evaluation set

Collect representative queries and mark which passages answer each one. These labels are relevance judgments. Keep test queries separate from the examples used to tune or train the system; these are held-out queries.

Include different wording, exact identifiers, negation, supported languages, and unanswerable questions. Sample queries to make evaluation manageable, but retain a realistic search corpus with competing documents. Removing difficult distractors can make a weak system look good.

Recall and precision: did we find the right passages?

Suppose a question has three relevant passages in the corpus. We return five results, of which two are relevant. In metric names, k means how many top results we inspect.

  • Recall@5 = 2/3: we found two of the three relevant passages. Recall asks, “How much of the useful material did we recover?”
  • Precision@5 = 2/5: two of our five results were relevant. Precision asks, “How much of what we returned was useful?”

A broad candidate search often emphasizes recall because a reranker cannot recover missing passages. The final list also needs precision so irrelevant material does not crowd out useful evidence.

MRR and nDCG: did the best answers appear early?

Mean reciprocal rank (MRR) focuses on the first relevant result. First place scores 1; second place scores 1/2; third place scores 1/3. Average these values across queries. If no relevant result appears within the evaluated list, that query contributes zero.

Normalized discounted cumulative gain (nDCG) can consider several relevant results and degrees of usefulness. For example, label a passage 0 for irrelevant, 1 for somewhat useful, and 2 for directly answering the question.

A common calculation is:

DCG@k⁡=∑i=1k2reli−1log⁡2(i+1) \operatorname{DCG@k}=\sum_{i=1}^{k}\frac{2^{rel_i}-1}{\log_2(i+1)}

For each position i, the numerator turns its relevance grade into a reward: grades 0, 1, and 2 become 0, 1, and 3. The denominator reduces the reward at later positions. Add the rewards to obtain discounted cumulative gain, or DCG.

Then divide by the DCG of the best possible top-k ordering of the judged items in the corpus:

nDCG@k⁡=DCG@k⁡ideal DCG@k⁡ \operatorname{nDCG@k}=\frac{\operatorname{DCG@k}}{\operatorname{ideal\ DCG@k}}

If grades [1, 2] appear in the first two positions, DCG is 1 + 3/log₂(3) ≈ 2.893. The ideal order [2, 1] gives 3 + 1/log₂(3) ≈ 3.631. Their ratio is about 0.797; the weaker ordering loses credit.

Specify the relevance scale, gain formula, and treatment of queries with no relevant passages. Incomplete labels also limit what these metrics can tell you. Sentence Transformers evaluation.

ANN recall answers a different question

We also need to check whether approximate search preserves the model's own best vector matches.

Suppose exact vector search returns ten neighbors, and ANN returns eight of those ten. ANN Recall@10 is 0.8. This measures agreement with exact search, not whether those neighbors answer the question. Even perfect ANN recall cannot rescue an embedding model that ranks irrelevant passages highly.

Use the failure to choose the next investigation:

What happened? What to investigate
Exact search finds useful passages; ANN misses them Search breadth, compression, and filtering
Both searches miss useful passages Embedding model, input formatting, chunk boundaries, or missing source material
Useful passages are retrieved but the answer is wrong Evidence selection, prompt/context assembly, and answer generation

What if the corpus has no answer?

A nonempty collection has a highest-scoring item even for an unrelated question. Returning the top item alone is not a relevance check.

Evaluate a rejection rule using answerable and unanswerable queries. It might use a similarity threshold, a reranker score, or a separate evidence check. Choose the threshold from measured results; a cosine value such as 0.8 is not universally meaningful.

If evidence is insufficient, say no supported answer was found or ask for clarification. Measure both mistakes: rejecting a query with a good answer and accepting one without support.

Watch for oddly generic matches

Sometimes the same passage appears near the top for many unrelated queries. The vector-space term hubness describes points that become neighbors of unusually many other points. Repeated boilerplate can also produce unhelpfully generic results, so inspect the text before diagnosing the geometry.

Anisotropy means vector directions are unevenly distributed, for example concentrated in a narrow region. This can make similarity scores less discriminating. Centering subtracts an average vector; whitening rescales and transforms the coordinates to reduce correlations. Both change the geometry, so test them rather than applying them automatically.

Batch work and reuse vectors

Sending one passage per request often wastes overhead. Batching sends several passages together. Stay within both the item limit and total token limit, and avoid batches too large for memory or acceptable response time.

Throughput measures how many items a system processes per second. Latency measures how long one request takes. Bigger batches can improve throughput while making individual items wait longer. Limit concurrent requests, retry temporary failures with increasing delays, and record completed work so a failed job can resume.

A cache stores results for reuse. If the input and embedding configuration are unchanged, reuse the vector instead of paying to compute it again.

Build the cache key from the complete input plus the model/revision, query/document role, instructions, preprocessing, output dimension, and normalization/precision settings. A hash can turn those fields into a compact lookup key; encode the fields unambiguously so different combinations do not accidentally form the same input string. Preserve punctuation and case when they carry meaning.

Include the authenticated tenant or permitted sharing scope for private inputs. A content hash is not an access-control decision or proof of anonymization. Keep raw inputs, vectors and cached query results under the same applicable access and retention rules. A revoked document must not reappear through an old result cache merely because its vector is unchanged.

Late chunking adds a dependency: a chunk's vector can change when its surrounding text changes. Include the encoded span and chunk boundaries in its cache identity. Recompute all affected chunks after an edit, potentially the whole span. Added contextual explanations are part of the embedding input too.

Keep access and deletion rules attached to the data

Keep source text, document/chunk IDs, versions, and access permissions alongside the vectors. Numeric form does not anonymize sensitive content.

Search only material the user may access, or apply a filtering strategy that guarantees unauthorized text never reaches the user, reranker, or answer model. If you retrieve ten passages and then discard nine for permissions, only one remains; the search strategy must account for that possibility.

When a document is deleted, remove its chunks from search and invalidate relevant cached results. Preserve source locations so an answer can cite where its evidence came from.

Handle document changes, usage drift, and model upgrades

These changes need different responses:

  1. Documents change: re-embed affected text and remove stale chunks. With late chunking, include affected neighboring context.
  2. Usage changes: users ask new kinds of questions or use new languages. This is usage drift. The old model can lose practical quality even if its weights never change; monitor with fresh evaluation queries.
  3. The representation changes: new model weights, prompts, pooling, or dimensions may produce vectors incompatible with the old ones. Record these settings with each index so they are not silently mixed.

For an incompatible model upgrade, use two indexes side by side. This is often called a blue-green deployment:

  1. Keep index A and its query encoder serving users.
  2. Build index B by encoding the corpus with the new model/configuration.
  3. Copy ongoing document updates and deletions into B while it is being built.
  4. Evaluate B. You can also run shadow queries: send copies of live queries to B for comparison without changing users' results.
  5. Switch query encoding and document search together to the new pair.
  6. Keep A available and current for a rollback if problems appear.

When comparing incompatible indexes, encode each query with the encoder appropriate to that index. Do not send a new-model query vector to an old-model document index or directly add their raw scores. If combining result lists is necessary, use an evaluated method such as rank fusion.

Not every model change forces a complete re-embedding. Documented shared spaces or supported vector transformations may allow reuse; an index rebuild can still be required. Small numerical variation is also different from an incompatible representation.

Estimate the full cost and response time

For 10 million chunks averaging 500 billed tokens:

10,000,000 × 500 = 5,000,000,000 input tokens
5,000 million-token units × $0.10 per unit = $500

The rate is illustrative, not a current vendor quote. Add repeated overlap, retries, query embeddings, storage, indexing, reranking, and the temporary second index during an upgrade. For self-hosting, estimate hardware cost and measured processing capacity instead.

Measure the whole request: query encoding + search + source-text fetch + reranking + network and queue delays. A 5 ms vector search does not imply a 5 ms service. p95 latency is the time at or below which 95% of measured requests finish; it reveals slow requests that an average can hide.

Interview exercise: migrate ten million passage vectors

Prompt: A private knowledge-search service has ten million passage vectors. A new embedding model improves difficult-query relevance and supports 512-dimensional output instead of the current 1,024 dimensions. Design a migration while documents, permissions and searches continue to change.

This is an interview scenario. The rates, targets and costs below are explicit assumptions, not a vendor benchmark or a claim about an actual customer deployment.

1. Clarify functional requirements

  1. Continue answering searches against permitted, current source documents.
  2. Build a new index using the new model's documented query/document formatting and 512-dimensional representation.
  3. Capture document additions, edits and deletions during the rebuild; preserve source IDs and versions.
  4. Compare relevance on representative queries, including hard negatives, exact identifiers, different languages and no-answer queries.
  5. Roll traffic to the new query encoder and index together, with a usable rollback path.
  6. Report progress, failures, missed updates and the expected completion time to operators.

Out of scope: training the embedding model, changing the answer-generation model and migrating unrelated application records. Keep these fixed initially so the experiment can identify the embedding change's effect.

2. Agree on non-functional requirements

  1. Availability: searches continue during backfill; one failed indexing worker cannot stop live serving.
  2. Latency: assume p95 end-to-end search below 500 ms at the agreed peak load, measured including query encoding, authorized text fetch and reranking.
  3. Freshness: assume ordinary content changes become searchable within 15 minutes. Measure backlog age rather than only processed item counts.
  4. Authorization: check current permissions before disclosing text to callers, rerankers or generators. An index's stale permission copy cannot authorize access.
  5. Quality: agree on relevance thresholds and important query slices before evaluating the new model. A good aggregate score cannot compensate for a severe access or language regression.
  6. Recoverability: resume interrupted work without losing updates, resurrecting deleted content or mixing incompatible representations.
  7. Cost: reserve live-query quota and bound migration workers, retries and temporary storage.

Permission revocation and content freshness are separate contracts. Define the authoritative permission read and its consistency boundary; do not promise instantaneous global revocation merely because a filter exists. If that authority is unavailable, fail closed for private source disclosure.

3. Start with the basic design and find its flaws

The smallest design exports documents, embeds them into index B, and changes the search configuration when the job finishes. Existing index A continues serving meanwhile.

Failure in that baseline Why it happens Repair and its cost
B misses an edit made after export A snapshot alone does not describe subsequent changes Capture a durable change stream; operate replay and lag monitoring
A deleted document reappears An old embedding job finishes after its deletion Version each document and retain deletion markers; add conditional writes
Queries use the wrong vector space Query encoder and index settings change independently Publish one immutable release bundle; pin it for each request
Cutover passes a superficial completeness check Equal row counts can hide missing IDs, wrong versions or duplicates Compare source identity/version coverage and inspect failures
Rollback returns stale content A is retained but no longer receives updates Keep both indexes current during a time-bounded rollback window
Backfill delays live searches Workers exhaust provider quota or shared storage capacity Separate budgets and reserve capacity for live requests

A deletion marker, often called a tombstone, records that a source version was removed even when no searchable text remains. Retain it long enough to reject older in-flight or replayed work; its exact retention follows the replay and deletion policy.

4. Build the detailed design

Architecture / visual model
flowchart TD S[Authoritative sources<br/>document IDs and versions] --> SNAP[Consistent snapshot<br/>record checkpoint W] S --> LOG[Durable change stream<br/>updates and deletion markers] SNAP --> WORK[Budgeted backfill workers<br/>new model and formatting] WORK --> APPLY[Conditional apply by document version] LOG --> REPLAY[Replay from W<br/>monitor lag and failures] REPLAY --> APPLY APPLY --> B[Index B<br/>new compatible passage vectors] LOG --> OLD[Existing index updater] OLD --> A[Index A<br/>current production vectors] U[Authenticated query] --> ROUTE[Pin release bundle<br/>query encoder plus index plus settings] ROUTE --> A ROUTE --> B A --> AUTH[Current authorization and source version check] B --> AUTH AUTH --> R[Rerank permitted text<br/>return evidence or reject] B --> GATE[Coverage and relevance checks<br/>shadow load and latency] GATE --> RELEASE[Controlled release change<br/>canary and rollback] RELEASE --> ROUTE
Read diagram source
flowchart TD
    S[Authoritative sources<br/>document IDs and versions] --> SNAP[Consistent snapshot<br/>record checkpoint W]
    S --> LOG[Durable change stream<br/>updates and deletion markers]
    SNAP --> WORK[Budgeted backfill workers<br/>new model and formatting]
    WORK --> APPLY[Conditional apply by document version]
    LOG --> REPLAY[Replay from W<br/>monitor lag and failures]
    REPLAY --> APPLY
    APPLY --> B[Index B<br/>new compatible passage vectors]
    LOG --> OLD[Existing index updater]
    OLD --> A[Index A<br/>current production vectors]
    U[Authenticated query] --> ROUTE[Pin release bundle<br/>query encoder plus index plus settings]
    ROUTE --> A
    ROUTE --> B
    A --> AUTH[Current authorization and source version check]
    B --> AUTH
    AUTH --> R[Rerank permitted text<br/>return evidence or reject]
    B --> GATE[Coverage and relevance checks<br/>shadow load and latency]
    GATE --> RELEASE[Controlled release change<br/>canary and rollback]
    RELEASE --> ROUTE

Ingestion and updates:

  1. Obtain a consistent source snapshot with a corresponding change-stream checkpoint W. If a connector cannot provide this pairing, design a connector-specific reconciliation procedure and measure its gaps before claiming lossless capture.
  2. Backfill the snapshot. Each job records tenant, source ID, source version, chunk identity, full embedding configuration and outcome. Make retries resume recorded work.
  3. Replay changes after W into B while backfill runs. Apply only versions newer than the stored version, using the database or indexing coordinator's atomic conditional update. A separately executed “read version, then write” has a race.
  4. Represent deletion as a versioned state too. An older backfill response cannot overwrite a newer deletion marker. If an index cannot enforce this condition atomically, serialize each document's updates through a coordinator until their writes finish. A version check separated from the index write is insufficient by itself; validate active versions again before using search evidence.
  5. For a changed document, stage the new chunk set and publish its document manifest only when complete. Search validates candidates against the active source/chunk version; incomplete or superseded chunks are excluded. An edit that reduces a document from ten chunks to six must retire the other four.
  6. Retry temporary failures with bounded attempts and backoff. Quarantine permanent failures for repair; do not call an index complete while required documents remain missing.

A checkpoint is a position in an ordered source stream, not necessarily a single global number for all connectors. Keep per-source or per-partition checkpoints where ordering is only local. See ingestion pipelines for connector recovery and reconciliation.

Evaluation and release:

  1. Compare B's document IDs and versions with an agreed source checkpoint. Confirm replay has caught up to that checkpoint and continues processing newer events.
  2. Evaluate exact retrieval first, then quantify losses from approximation, filters and quantization. Keep labels and the candidate corpus consistent across model comparisons.
  3. Send permission-scoped shadow queries through B's own encoder. Measure relevance, no-answer decisions, p95 latency, query cost and queue behavior at expected load. Shadowing private requests must preserve their access and retention rules.
  4. Create a release bundle containing model revision, input formatting, output dimension, normalization, index ID and compatible reranker configuration. Each request keeps the bundle it started with, including during a rollout.
  5. Shift a small, defined share of traffic to B. Expand only when the agreed gates hold. Stop or roll back on significant quality, latency, cost or correctness regressions.
  6. Continue updating A during the rollback window. After the window, drain requests pinned to A and remove obsolete vectors, caches and retained source copies according to policy.

An atomic routing change does not make every distributed component switch simultaneously. Pinning a complete bundle makes both old and new requests internally compatible while they overlap.

5. Estimate time, storage and full cost

Assume ten million chunks averaging 500 billed input tokens, 12 workers each sustaining 20 chunks/second, and 70% planned worker utilization.

Worker capacity = 12 × 20 × 0.70 = 168 chunks/second
Provider quota = 6,000,000 tokens/minute
Reserve 20% for live work: backfill gets 4,800,000 tokens/minute
Quota-limited backfill = 4,800,000 / 500 / 60 = 160 chunks/second
Effective backfill rate = min(168, 160) = 160 chunks/second
With 5% additional billed attempts: 10,500,000 / 160 / 3,600 ≈ 18.23 hours

This is a capacity estimate, not a completion guarantee. Add index construction, change replay, variable input lengths, validation and incidents. Some failed attempts consume quota without billing, or billing without useful output; use the provider's actual contract when budgeting.

Raw-vector storage Calculation Decimal GB
Old A, 1,024 dimensions, FP32 10 million × 1,024 × 4 bytes 40.96
New B, 512 dimensions, FP32 10 million × 512 × 4 bytes 20.48
Both indexes, one copy 40.96 + 20.48 61.44
Both indexes, two copies each 61.44 × 2 122.88

These numbers exclude index graphs/codebooks, text, metadata, checkpoints, backups and spare capacity. Shorter output vectors do not automatically reduce the encoder's inference time or per-token API price.

Migration cost assumption Calculation Cost
Embedding calls including 5% extra attempts 5.25 billion tokens × $0.10/million $525
Incremental temporary index storage Assumed bill for the migration window $80
Worker infrastructure Assumed total for the job $300
Engineering and operational preparation 24 hours × $80/hour $1,920
Evaluation and review 6 hours × $60/hour $360
Total migration cost Sum of the rows $3,185

The $80 and $300 are budget assumptions, not prices derived from a particular cloud SKU. Existing live-service costs are excluded because this is an incremental migration comparison. For a complete service budget, add recurring query encoding, indexes, reranking, network and operations.

Decision: the raw-vector footprint halves after A is retired, but savings depend on what proportion of the actual bill is vector storage. If the measured recurring benefit is $200/month, the $3,185 migration pays back in about 15.9 months on cost alone. A larger measured relevance or latency benefit may justify it sooner; an unmeasured leaderboard improvement does not.

6. Close the interview

“I would keep the current encoder/index pair live, build a versioned replacement, and reconcile concurrent updates and deletions before release. The core correctness risks are incompatible vector spaces, stale permissions and backfill races. I would gate the rollout on permitted-source coverage, relevance and end-to-end latency, and keep the old pair current until rollback is no longer needed. Our estimate is roughly 18.23 hours of embedding work and $3,185 incremental cost under the stated assumptions; actual quota behavior, index overhead and evaluation results decide the schedule and whether we proceed.”

Interview Recall

Use these cards for quick revision after the worked explanations.

Recall cue Answer to build from
What does an embedding preserve? Relationships learned for a task, not a universal measure of truth
Same dimensions, same space? No; representations must be explicitly compatible
Cosine versus dot product? Cosine removes length; dot product includes it; unit vectors make exact rankings equivalent
Two kinds of recall? ANN recall compares with exact neighbors; task recall compares with relevant evidence
Can reranking repair missing candidates? No; it only orders candidates it receives
Migration invariant? A request uses one compatible encoder/index bundle and current authorization
Best cost metric? Full cost at the required relevance, latency and freshness, not dimensions alone

Manager practice: explain the search pipeline in two minutes, then defend the migration's quality gate, capacity reservation and rollback window. Distinguish a model upgrade from changing user needs: query drift can reduce quality without any model change. Connect diagnosis to RAG evaluation before funding a new model or index.

Developed interview questions and answers

1. Two texts have cosine similarity 0.9. Are they equivalent?

No. The score describes their learned representation under a model and preprocessing configuration. It is not a calibrated 90% probability of equivalence. Contradictory statements can share a topic and score highly. Inspect the task: search relevance, paraphrase detection and entailment require different evidence and evaluation.

Follow-up: does cosine zero prove unrelated meanings? No; it means orthogonal vectors in that representation.

2. How do embedding models learn semantic similarity?

A common approach supplies positive pairs and negative examples, then optimizes a contrastive loss so positives score above negatives. Retrieval positives need to be useful answers, not necessarily paraphrases. A hard negative about the correct product but wrong procedure teaches a useful distinction. A relevant passage mislabeled as negative teaches the wrong one. Training objectives differ, so inspect the model card and test the actual domain.

Follow-up: are all other examples in a batch safe negatives? No; duplicates and multiple correct answers can create false negatives.

3. Both models return 1,024 dimensions. Can their vectors share one index?

Dimension agreement only satisfies a shape requirement. Coordinates from independently trained models need not have comparable meaning. Use separate compatible spaces unless the provider explicitly supports a shared representation with the chosen settings. A documented query/document encoder pair can differ internally and still be compatible; arbitrary equal-length vectors cannot establish that.

4. When would you choose ColBERT-style late interaction?

Consider it when single-vector compression loses important token-level distinctions and measured relevance gains justify extra storage and search work. Compare against a bi-encoder baseline on difficult, multi-part queries under the same latency and memory limits. A cross-encoder instead jointly processes each query-candidate pair; it is a useful reranker but a different computational architecture.

Follow-up: can a reranker repair a relevant document absent from the candidate set? No.

5. How do you choose embedding dimensions?

Start with supported dimensions and evaluate relevance, latency and total storage. Use trained shortening such as Matryoshka representations where documented; arbitrary truncation can damage rankings. Ten million 1,024-dimensional FP32 vectors require 40.96 decimal GB before indexes, text and replication. Halving dimensions halves those raw bytes, not necessarily the whole bill or embedding inference cost.

6. Exact search works; ANN search misses useful passages. What changes first?

Keep the representation fixed and investigate approximation: search breadth, compression, partitioning and filters. For HNSW, increasing exploration can improve recall at extra work; for IVF, probing more lists can help. Compare latency and ANN recall on the same permission-filtered workload. If exact search is also poor, index tuning alone cannot repair the model, chunking or missing evidence.

7. Why does dense retrieval miss exact error codes?

An embedding model may emphasize broad topic similarity over a tiny identifier change. Add a lexical or exact-field path with an analyzer that preserves meaningful characters, then evaluate hybrid candidate fusion and reranking. Copying a raw cosine score into a sum with BM25 assumes comparable scales; use calibrated scoring or an evaluated rank-fusion method.

8. Can a long-context embedding model replace chunking?

A larger input limit allows more context but does not guarantee one vector can retain every small fact for retrieval. Large chunks also increase downstream evidence cost and may mix permissions or versions. Compare coherent chunk boundaries, contextual enrichment and late chunking. Late chunking pools token representations after encoding a larger span; it needs access to those contextual token representations, not just an API returning one final vector.

9. What would you check when nearest neighbors look plausible but useless?

First confirm that the permitted, current corpus contains the answer and parsing preserved it. Compare exact with approximate retrieval, check encoder compatibility and role formatting, then inspect hard-negative errors and labels. Related-topic matches may benefit from better training, hybrid retrieval or reranking. Missing evidence requires source repair or an honest no-answer result. Do not spend days tuning index parameters for a data problem.

10. How do you change embedding models without corrupting retrieval?

Version formatting, model revision, dimensions, normalization and index as one release. Build the new corpus representation while preserving ongoing updates and deletions. Evaluate each index with its matching query encoder, then canary a complete bundle. Keep the old pair current for rollback. Re-embedding only queries is safe only when the new query space is documented as compatible with existing document vectors.

11. A deleted document reappears during backfill. What went wrong?

A stale job applied an older document after the deletion. Use a source version and deletion marker, and atomically reject older updates. Merely checking a version before a later unconditional write leaves a race. Also reject inactive document/chunk versions when fetching evidence, and invalidate cached results. Do not resolve the incident by deleting the row once while stale workers can write it again.

12. Can we share embedding caches across tenants using a text hash?

Only when sharing is explicitly permitted and the data contract supports it. A hash is neither authorization nor anonymization. For private data, key by authenticated scope and the full representation configuration, including context dependencies for late chunking. Recheck access before returning cached source text, and apply retention/deletion rules to vectors and results as well as raw inputs.

13. Why can a benchmark winner lose on our application?

Tasks, languages, negative examples and corpus composition can differ. A broad average can conceal poor retrieval on the application's most important slices. Evaluate real query categories with a realistic corpus and agreed labels, then compare whole-service latency and cost. Keep a holdout set separate from tuning; repeated threshold selection on the test set weakens the evidence.

14. How would you reduce vector storage without silently reducing quality?

Establish full-precision exact retrieval as a reference, then separately test supported shorter dimensions, scalar/PQ/binary compression and ANN settings. Report relevance as well as agreement with the reference. A compressed shortlist can be rescored with more precise vectors, but storing both representations reduces net savings. Include graph/codebook and replica overhead before claiming a percentage reduction in the bill.

15. Should a system always return the top result?

No. A nonempty corpus has a highest score even for an unrelated query. Calibrate an evidence or rejection rule on both answerable and unanswerable queries, report false acceptance and false rejection, and inspect important slices. Thresholds are model- and workload-specific. If evidence is insufficient, return a clear no-answer result or request clarification instead of treating rank one as proof.

Final notes

Decision Evidence needed Interview trap to avoid
Embedding model Task relevance, language coverage, input handling and license “Newest” or largest must win
Similarity metric Model's documented metric and normalization Cosine is a probability
Chunking Complete evidence units, query performance and access boundaries Maximum input length is the best chunk size
Search index Exact baseline, ANN recall, filtered latency and memory Perfect ANN recall means useful results
Model migration Version coverage, compatible routing and current permissions Same dimensions make old and new vectors interchangeable
Cost reduction Full recurring and migration costs at acceptable quality Raw-vector savings equal service savings

Interview tip: define the term first, show one small numeric example, then discuss its system consequence. For embeddings: learned representation → compatible comparison → useful evidence → measured tradeoff. Keep correctness, authorization and relevance as separate checks.

References

Previous: Transformer Architecture | Next: Inference Pipeline

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Transformer Architecture: From Recurrent Memory to Modern AI
NEXT LESSONInference Pipeline →

Explore the diagram