Multimodal retrieval-augmented generation retrieves evidence involving more than one type of information, such as text, images, audio or video, and uses that evidence to produce an answer. A table's structure and a chart's visual encoding can matter even when their labels are text. The retrieval model, extraction model and answer generator may be different components.
The engineering question is which information the task needs and how each representation preserves it. There is no universal percentage of enterprise knowledge contained in images, or one architecture that is best for every document collection.
Identify what a text-only path can lose
| Source | Information at risk | Evidence to preserve |
|---|---|---|
| Table | Header hierarchy, units, footnotes and row relationships | Structured cells plus source location |
| Chart | Axes, scale, legend, uncertainty and plotted values | Original image and underlying data when available |
| Architecture diagram | Arrow direction, labels and component grouping | Diagram regions and typed relationships |
| Scanned page | Text absent from the PDF's text layer | Page image and validated OCR/extraction |
| Audio | Words, speaker turns, timing and relevant sounds | Audio segments, timestamps and transcript provenance |
| Video | Temporal events and alignment between sound and frames | Time ranges, frames/clips and aligned transcripts |
A good text parser can preserve many tables and headings. Conversely, a vision model can misread a small number or invent a relationship. Evaluate the actual source and task instead of assuming text extraction always fails or vision always succeeds.
Compare the representation choices
| Design | How retrieval works | Benefit | Main limitation |
|---|---|---|---|
| Describe and index | Extract text or generate descriptions, then use text retrieval | Reuses lexical/dense infrastructure | Descriptions may omit or invent details |
| Shared multimodal embedding space | Encode supported modalities with a compatible model | Cross-modal similarity search | Similarity does not ensure exact reading or arithmetic |
| Separate retrieval branches | Use appropriate text, visual and structured retrieval, then fuse | Tune each evidence path | More indexes, routing and evaluation |
| Page-image late interaction | Compare query token vectors with page representations | Preserves visual layout for retrieval | Larger representations and document-image processing |
These choices can be combined. A report search service may use extracted text for exact identifiers, a visual retriever for diagrams, and structured queries for validated tables.
Current model families and their roles
CLIP-style dual encoders and SigLIP-family models align image and text representations. SigLIP 2 adds training techniques for semantic understanding, localization and dense features; model suitability still depends on the target documents and queries. A natural-image retrieval result does not establish accuracy on a dense financial table. SigLIP 2 paper.
Examples of current managed embedding options include Cohere Embed v4, Voyage Multimodal 3.5 and Gemini Embedding 2. Their supported media, input composition, dimensions and limits differ. Gemini Embedding 001 is text-only; it must not be described as an image embedding model. Gemini Embedding 2 uses a different representation space, so migration requires re-embedding. Cohere, Voyage, Google.
A vision-language generator can read selected images or extract structured information, but an embedding endpoint does not itself generate an answer. Validate candidate generators using small text, tables, cross-page references, refusal behavior, output schemas, media limits and data-handling requirements. Avoid unsourced “excellent/good” vendor scorecards.
Visual document retrieval with late interaction
ColPali uses a vision-language backbone to produce multi-vector page representations and scores them against query representations with late interaction. This can retrieve visually rich pages without requiring an OCR pipeline to construct the retrieval representation. Its ViDoRe benchmark evaluates page retrieval, not complete application correctness. ColPali paper.
ColQwen2.5 variants use Qwen2.5-VL; they should not be mislabeled as using Qwen2-VL. Model cards define the actual backbone, processor, image handling and licensing requirements. Review both backbone and adapter terms. ColQwen2.5 model card.
“No OCR required for this retriever” does not mean no extraction is useful anywhere. Searchable citations, exact identifiers, accessibility, redaction, numerical checks and structured analysis may still require parsed text or data. Nor does a page model necessarily emit exactly 1,024 vectors for every configuration; resolution and processor behavior matter.
See late interaction and MaxSim for the scoring and storage tradeoffs.
Design a report question-answering service
Use an illustrative collection of product research reports containing prose, tables and charts. The service must explain reported measurements and compare periods without inventing values.
Functional requirements:
- Ingest supported document types and retain versioned originals.
- Search permitted text, tables and visual evidence for the question.
- Answer with page/region citations and show the source evidence.
- Perform requested calculations using validated values and explicit units.
- Clarify ambiguous periods or report insufficient evidence.
Non-functional requirements:
- A stated source freshness target and complete-request latency budget.
- Access control on raw files, derivatives, snippets and citations.
- Observable extraction and retrieval errors, with retry-safe ingestion.
- Measured quality by document type, language and image readability.
- Bounded storage, model cost, image resolution and query context.
Start with a parser and text retrieval on representative reports. Inspect actual failures. Add table structure or visual retrieval where it improves those failures; do not require three indexes simply because the input is a PDF.
Evolve the architecture from that baseline
Read diagram source
flowchart TD
S[Versioned source files and access rules] --> P[Classify and parse supported content]
P --> T[Text and heading records]
P --> V[Page images and selected regions]
P --> B[Validated table records]
T --> TI[Lexical and dense indexes]
V --> VI[Visual representation index]
B --> BI[Structured store and searchable descriptions]
Q[Query and authenticated scope] --> R[Choose evidence paths]
TI --> C[Retrieve eligible candidates]
VI --> C
BI --> C
R --> C
C --> F[Fuse identities and assess evidence coverage]
F --> A[Fetch permitted source regions and data]
A --> N[Validate values and calculate when needed]
N --> G[Generate answer with source references]
G --> U[Validate citations and display answer]
Parsing tools such as Docling can preserve document structure and support multiple source types. Choose the parser and OCR/extraction configuration from measured errors. Rendering at a fixed 300 DPI for every page is not automatically the best cost/quality choice. Preserve the coordinate system and transformation used for any resize or crop. Docling documentation.
Keep raw assets in private object storage and derivative records in suitable stores. Authorize access before sending evidence to a model and before returning an asset URL. A public URL to a private page image defeats permission filtering in the retriever.
Preserve a useful evidence contract
For each derived item, record source identity, revision, modality, location, extraction/model versions, permissions and validation status. For example:
{
"evidence_id": "report-82:r3:p7:table2",
"source_id": "report-82",
"source_revision": "r3",
"page_number": 7,
"region_xyxy_normalized": [0.12, 0.25, 0.88, 0.72],
"kind": "table",
"units": "milliseconds",
"validation_status": "checked_against_source",
"derived_from": ["report-82:r3:p7"]
}
Here coordinates are normalized to page width and height, with origin at the top left; page numbers are one-based. Define conventions explicitly. The status field is set by the validation process, not accepted merely because a model wrote it.
Two hits from a chart image and its generated caption are two representations of one source, not independent corroboration. Deduplicate and preserve that relationship when fusing candidates.
Tables: preserve semantics, not an absolute chunk rule
Small tables can remain whole. Large tables may exceed model or retrieval budgets and need row groups or structured queries. When splitting, repeat relevant headers, preserve row identifiers, units, footnotes and links to the full table. A blanket “never split tables” rule fails on thousands of rows.
For exact sums, filters or joins, prefer validated structured data and deterministic computation over asking a language model to mentally aggregate a large Markdown table. Extraction still needs checks: a syntactically valid number in the wrong column is wrong data.
Suppose a validated table reports p95 latency of 240 ms before a change and 180 ms afterward. The reduction is (240 − 180) / 240 = 25%. Confirm that both measurements use comparable workloads, units and percentile definitions before making that comparison. A chart caption saying “improved performance” does not supply missing measurement conditions.
Charts, diagrams, audio and video
For a chart, retain axes, units, scale, series legend, error bars and applicable footnotes. When exact numbers matter, use the underlying dataset if available. Values visually estimated from a chart should be labeled approximate; logarithmic axes and truncated baselines can change interpretation.
For a diagram, verify arrow direction and labels. Spatial proximity is not proof that two services communicate. Extracted nodes and edges should retain the source regions needed to inspect the claim.
For audio, align transcript segments with timestamps and preserve speaker uncertainty. Transcription errors can change names, negation or numbers. Questions about a sound may require audio evidence rather than transcript text alone.
For video, frame sampling can miss brief events. Use time-aligned transcript, frames and clips according to the task, with bounded temporal expansion when a result needs surrounding context. Cite the time range used. A thumbnail and an unrelated transcript segment should not become a fabricated combined event.
Cross-modal retrieval and context assembly
A question can need a chart on one page and a table on another. Those lookups may be independent and run in parallel; cross-modal does not inherently mean multi-hop. A dependency exists when one result determines the next lookup.
Route by required evidence rather than imposing fixed quotas such as “two images for every question.” Quotas can help a tested diversity policy, but irrelevant images consume context and may distract the generator. Track coverage of the actual requested facts.
Rank fusion can combine independent rankings without assuming their raw scores share a scale. A text-only cross-encoder cannot directly assess raw images; use a compatible multimodal reranker or an explicitly evaluated text representation. Preserve original images for verification when a caption is insufficient.
If evidence remains incomplete, issue a bounded follow-up lookup, clarify the question or abstain. Do not force a visual model to infer an unreadable value.
Budget and test the design
Assume one million pages, 1,024 vectors per page, 128 dimensions and two bytes per component. The raw vector payload is 1,000,000 × 1,024 × 128 × 2 = 262.144 GB in decimal units. These are illustrative configuration assumptions. Add image bytes, indexes, text, metadata, replicas and rebuild headroom.
Quantization can reduce representation size, but “32× smaller” applies only to float32 components reduced to one bit before overhead. It is not automatically the reduction in total storage or a guarantee of a small recall loss. Test tiny text, rare identifiers and subtle visual distinctions after compression.
Ingestion cost recurs on changed sources or processors. Measure render, extraction, embedding and validation costs separately, including retries. For query latency, measure the slowest parallel branch plus fusion, asset fetch, generation and queues; do not add all branch durations as though they run serially.
| Failure test | Expected behavior |
|---|---|
| Header or unit missing | Reject uncertain extraction or load source context |
| Unreadable chart value | Report uncertainty; do not invent precision |
| Same page appears in three branches | Merge identity without counting three independent sources |
| Permission revoked after indexing | Deny evidence and image access under the current policy |
| Model reads instructions inside an image | Treat them as untrusted source content |
| Correct page, wrong table cell | Fail answer validation despite page-level retrieval success |
Interview practice
Q1: When is visual retrieval worth adding?
When reviewed failures show that layout, images or visual relationships are needed and the text path loses them. I would compare a better parser, describe-and-index and direct visual retrieval on those query slices before choosing the extra infrastructure.
Q2: Does ColPali eliminate every need for OCR?
It can construct visual retrieval representations without OCR. The application may still need extracted text for exact search, accessibility, citations, redaction or structured validation. Retrieval and evidence consumption are different stages.
Q3: How do you answer an exact numeric question from a table?
Resolve the correct source/version, preserve headers and units, validate the selected cells and compute using deterministic code when arithmetic is needed. Cite the relevant location. If extraction is uncertain, return uncertainty or obtain a clearer source.
Q4: Can a shared vector space replace all separate indexes?
It can simplify cross-modal similarity search for supported inputs. Exact filtering, structured aggregation, permissions and modality-specific quality may still justify separate paths. Sharing a vector space does not make every task a nearest-neighbor problem.
Q5: How do you combine chart and table evidence?
Retrieve the required facts, verify compatible periods and units, preserve provenance, then assemble a bounded evidence packet. Use parallel lookups when independent and additional retrieval only when needed. Fixed modality quotas do not guarantee coverage.
Q6: What is different about evaluating this system?
I check extraction and location accuracy, modality-specific retrieval, numerical fidelity, cross-source alignment and citation support. A correct page hit may still produce a wrong answer from the wrong cell, axis or time range.
Q7: Where do access checks belong?
On source ingestion and policy assignment, retrieval, derivative resolution, model submission and asset delivery. Derived captions, crops and embeddings retain the source's access obligations; an unrestricted image URL is a disclosure path.
Final notes
Recall card: Preserve modality-specific meaning → retrieve the required evidence → validate location and values → answer with inspectable sources. The strongest design adds visual or temporal processing where it solves a demonstrated information loss.