Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Chunking Strategies

By Anup Rai5 min readReviewed September 2026

Chunking divides source material into units that can be indexed, retrieved and supplied as evidence. The search unit and the unit shown to the answering model do not have to be identical. A small matching passage may lead to a larger section containing its definitions and exceptions.

The goal is to preserve answering evidence while controlling retrieval selectivity, context size, ingestion cost and update behavior. There is no universally correct chunk length or overlap percentage.

Start with the evidence a question needs

Consider a policy section:

Returns — refurbished devices
Unopened devices may be returned within 30 days of delivery.
Opened devices may be returned only within 14 days of delivery.
These limits do not replace separately stated regional requirements.

For “Can an opened refurbished device be returned after 20 days?”, the answer needs the product category, opened-device rule, time basis and any applicable regional exception. Indexing only “within 30 days” produces a searchable but misleading fragment.

First inspect parsing. Overlap cannot repair a parser that detached a table cell from its header or changed a decimal in OCR.

Compare strategies by their failure modes

Strategy How it forms units Useful property Main limitation
Fixed token windows Size and overlap Simple, reproducible baseline Can split meaning and structure
Recursive splitting Tries paragraphs, sentences and smaller boundaries Respects common text structure Punctuation is not always a semantic boundary
Structure-aware Uses headings, sections, tables or syntax Preserves document organization Depends on parsing quality
Semantic segmentation Detects topic changes with learned representations Can adapt to prose organization Adds inference cost and threshold sensitivity
Parent-child retrieval Searches children and expands selected parents Selective matching with surrounding context Parent expansion can fill the context budget
Late chunking Contextualizes tokens before pooling chunk representations Retains surrounding context in embeddings Requires suitable model access and bounded document context

Small chunks do not inherently have higher precision or lower search latency. More chunks mean more index entries, and a short fragment can be ambiguous. Larger chunks can preserve meaning but mix unrelated topics and consume generation tokens. Measure these effects together.

Design a parent-child pipeline

Architecture / visual model
flowchart LR D[Parsed section with source metadata] --> P[Store parent section] D --> C[Create searchable child passages] C --> I[Index children with parent IDs] Q[Question] --> I I --> H[Select permitted matching children] H --> R[Load and deduplicate relevant parents] P --> R R --> B[Pack evidence within context budget]
Read diagram source
flowchart LR
    D[Parsed section with source metadata] --> P[Store parent section]
    D --> C[Create searchable child passages]
    C --> I[Index children with parent IDs]
    Q[Question] --> I
    I --> H[Select permitted matching children]
    H --> R[Load and deduplicate relevant parents]
    P --> R
    R --> B[Pack evidence within context budget]
  1. Preserve source ID, version, section path and permissions.
  2. Create children whose boundaries retain useful claims.
  3. Store the relationship from child to parent.
  4. Retrieve permitted children using the chosen search method.
  5. Expand only as much parent context as the question needs.
  6. Deduplicate overlapping parents and retain precise citation offsets.

Both children and parents may be indexed if evaluation supports it. “Only index children” is a design choice, not the definition of hierarchical retrieval. A parent must not contain material the user is forbidden to read simply because one child is permitted.

Calculate overlap and expansion cost

For a 10,000-token document, windows of 500 tokens and overlap of 100 advance by 400 tokens. The count is ceil((10,000 − 500) / 400) + 1 = 25 windows. With a shorter final window, the total indexed text is 10,000 + 24 × 100 = 12,400 tokens, before repeated headings or metadata.

With no overlap, there would be 20 full windows and 10,000 indexed tokens. The additional 2,400 tokens may improve boundary coverage, but also increase embedding work and duplicate candidates.

Suppose five matching children point to the same 2,000-token parent. Loading that parent once uses 2,000 evidence tokens; blindly loading it for every child uses 10,000. Deduplication is part of context packing, not an optional cosmetic step.

Handle content types deliberately

Content Preserve When it is too large
Code Symbol path, signature, relevant imports and source lines Split by meaningful blocks with references to the enclosing symbol
Tables Column headers, units, row identifiers and footnotes Retrieve relevant row groups plus headers and linked source table
PDFs Reading order, page coordinates, captions and image references Use layout-aware regions or page retrieval with targeted follow-up
Lists and procedures Step order, preconditions and exceptions Include dependencies when expanding a matched step

“Never split a function” is impractical for generated or very large functions. Preserve structure and references while respecting model limits. A table summary can help retrieval, but the original table remains the evidence for exact values.

ColPali is a visual document retriever that embeds page images into multiple vectors and uses late interaction. It is not a generic PDF parser that guarantees correct text reading order. ColPali paper.

Distinguish context-enrichment techniques

Contextual prepending adds a short description of a chunk's role in its document before indexing. Anthropic's contextual retrieval applies this idea to embeddings and BM25. Verify the generated description; it can introduce unsupported context. Contextual retrieval.

Late chunking first runs a compatible long-context embedding model over a larger text, then pools token representations over chunk spans. It does not mean retrieving a whole document and splitting it afterward. The embedding has seen surrounding content even though the stored representation is chunk-level. Late chunking paper.

Interview practice

Q1: How do you select chunk size?

Start with document structure and representative questions. Compare evidence coverage, answer quality, retrieved-token volume and ingestion cost across a small set of policies. Inspect boundary failures rather than assuming 500 tokens is universally suitable.

Q2: What does overlap solve?

It reduces loss around boundaries by repeating nearby content. It does not restore missing headings, fix OCR or guarantee that a distant exception remains attached. It also increases indexing work and duplicate results.

Q3: Why retrieve children but return parents?

Children can provide more selective matches; a parent can supply definitions and qualifications. The tradeoff is larger evidence payloads. Deduplicate parents, enforce their permissions and expand only the relevant context.

Q4: Would you always choose semantic chunking?

No. It introduces model and threshold dependencies and may not match the task's evidence boundaries. A clear section structure or a measured fixed-window baseline can be sufficient. Compare its actual benefit before adding ingestion complexity.

Q5: How do chunking changes affect updates?

Boundary changes can alter many chunk IDs and embeddings. Keep stable source/version lineage, make ingestion idempotent and delete stale descendants. Evaluate the replacement index before cutover and preserve a rollback path.

Q6: Is late chunking the same as ColBERT?

No. Late chunking contextualizes tokens before pooling them into chunk embeddings. ColBERT retains multiple token representations and combines query-document similarities at scoring time. They address different stages and have different storage costs.

Final notes

Recall card: Preserve structure → select evidence units → measure boundaries → control expansion. Chunking is a retrieval and evidence-design decision, not a magic token count.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← RAG Fundamentals
NEXT LESSONEmbedding Models →

Explore the diagram