Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

OCR and layout analysis: turn documents into reliable evidence

By Anup Rai8 min readReviewed September 2026

Optical character recognition (OCR) converts images of text into machine-readable characters. Layout analysis identifies document regions and their relationships: paragraphs, headings, columns, tables, figures, captions, and reading order. Information extraction then maps that content into fields or records. These are related tasks, not synonyms.

A page can have perfect character recognition and still become useless when a parser interleaves two columns or assigns a table value to the wrong heading. Conversely, a plausible summary can hide omitted or incorrectly recognized source text.

Start with the document, not the model

Input First path to evaluate What to verify
Born-digital PDF with usable text Extract text and positional information directly Reading order, encoding, missing glyphs and image-only regions
Scanned page OCR with layout analysis Language support, resolution, orientation and recognition errors
Mixed PDF Route pages or regions according to content Avoid duplicate text from an existing OCR layer
DOCX, HTML or spreadsheet Parse native structure where available Headers, tables, formulas versus displayed values
Complex chart, diagram or unusual layout Targeted vision/layout interpretation Source grounding, labels, units and omissions

Rasterizing every document is not a universal standard. It can discard useful embedded text and structure. A PDF text layer can also be wrong or poorly ordered, so its presence alone does not establish quality.

Specialized OCR, document-layout models and vision-language models remain useful alternatives or components of a hybrid system. Azure Document Intelligence extracts structural elements as well as text. Docling's document representation preserves content items and a hierarchy that encodes reading order. A generic vision model is not the only way to understand layout.

The four outputs to keep separate

  1. Source artifact: the original bytes, document version and page images when generated.
  2. Transcription: recognized text tied to page/region evidence.
  3. Structure: headings, reading order, table cells and relationships.
  4. Interpretation: summaries, explanations and extracted business fields.

Never silently replace transcription with a model's corrected paraphrase. If the scan says an ambiguous O or 0, preserve uncertainty or request review when the distinction matters. A generated explanation belongs in a separate field.

An application-owned element record might look like this:

{
  "documentId": "doc-42",
  "sourceRevision": "sha256:...",
  "page": 7,
  "elementId": "p7-table2-r3-c2",
  "kind": "table_cell",
  "text": "120 ms",
  "row": 3,
  "column": 2,
  "columnHeader": "p95 latency",
  "coordinates": {
    "space": "normalized_top_left",
    "left": 0.55, "top": 0.40,
    "width": 0.18, "height": 0.04
  },
  "status": "needs_review",
  "extractorRevision": "layout-pipeline-v3"
}

This is an illustrative internal schema, not a provider's exact response. Record confidence only when its source and meaning are known; do not invent a confidence number from the model's prose.

Reading order: position is not sequence

For a two-column page, the logical order may be:

Title across both columns
┌─────────────────┬─────────────────┐
│ Left paragraph 1│ Right paragraph 3│
│ Left paragraph 2│ Right paragraph 4│
└─────────────────┴─────────────────┘
Full-width table and caption

Sorting all words by vertical position would interleave the columns. Detect regions, determine their ordering, and then order content within each region. A full-width figure or table can interrupt the column flow. Sidebars, footnotes and captions require explicit relationships.

Useful checks include whether sentences jump between columns, repeated headers enter every retrieval chunk, footnotes attach to the wrong paragraph, and a heading is incorrectly grouped with the preceding section. Both specialized pipelines and vision models can make these mistakes; evaluate them on representative pages.

Tables need more than Markdown pipes

Store row/column indices, header roles, cell spans, units and source regions. Merged cells and multirow headers may require structured JSON or HTML for faithful representation. Azure's documented layout output uses HTML tables for such cases; plain Markdown tables cannot represent every span cleanly.

A retrieval chunk containing 120 without its p95 latency header and ms unit loses meaning. Keep the needed headers with the row, preserve table identity, and cite the source page. For tables continued across pages, verify matching structure and continuation evidence before joining them.

Coordinates and preprocessing

Providers use different coordinate conventions. Textract bounding boxes use ratios of page width and height with a top-left origin. Other outputs can use PDF points, pixels or polygons.

For a normalized box with left=0.25, top=0.10, width=0.50, height=0.05 on a 2,000 × 3,000 pixel image:

x = 500 px; y = 300 px; width = 1,000 px; height = 150 px

If the image was cropped, rotated or resized, preserve that transform to map the extracted region back to the original page. A plausible bounding box is not sufficient evidence that a redaction covers every sensitive glyph.

Tesseract's quality guidance discusses preprocessing and segmentation. Evaluate deskewing, orientation correction, contrast, borders and noise handling against the input type. Aggressive cleanup can erase decimal points, faint characters or handwritten marks. Keep the source and compare before/after quality.

Vision models can also misread small or rotated text and can struggle with precise spatial localization. These are documented vision limitations, not problems automatically solved by “visual attention.” Neither OCR nor a vision model promises 100% character accuracy.

Interview design: ingest a 500-page technical handbook

Functional requirements

  1. Accept an authorized upload and return a durable ingestion job ID.
  2. Extract text, headings, tables, figures and page-level provenance.
  3. Identify failed or uncertain pages and support selective reprocessing/review.
  4. Create searchable chunks with stable source citations.
  5. Support replacement and deletion of a document and its derived artifacts.

Non-functional requirements

  1. Enforce file/page/resource limits and isolate document parsing.
  2. Preserve account access controls throughout extraction and retrieval.
  3. Bound concurrency, retries, provider spend and queue delay.
  4. Make partial progress recoverable without reprocessing every page.
  5. Measure extraction quality by document type and language.

The initial design can use a document store, job queue and one extraction worker. Persist page results before building an index. Do not make a 500-page upload depend on one long-running browser request.

Architecture / visual model
flowchart TD U[Authorized upload] --> S[Immutable source and job manifest] S --> P[Inspect pages and native text] P --> Q[Bounded page or region queue] Q --> R{Extraction route} R --> N[Native text and structure] R --> O[OCR and layout] R --> V[Targeted vision processing] N --> E[Versioned evidence records] O --> E V --> E E --> C[Quality and completeness checks] C -->|Uncertain| H[Review or selective retry] H --> E C -->|Accepted| A[Assemble reading order and cross-page structure] A --> K[Structure-aware chunks with access metadata] K --> I[Search and vector indexes]
Read diagram source
flowchart TD
    U[Authorized upload] --> S[Immutable source and job manifest]
    S --> P[Inspect pages and native text]
    P --> Q[Bounded page or region queue]
    Q --> R{Extraction route}
    R --> N[Native text and structure]
    R --> O[OCR and layout]
    R --> V[Targeted vision processing]
    N --> E[Versioned evidence records]
    O --> E
    V --> E
    E --> C[Quality and completeness checks]
    C -->|Uncertain| H[Review or selective retry]
    H --> E
    C -->|Accepted| A[Assemble reading order and cross-page structure]
    A --> K[Structure-aware chunks with access metadata]
    K --> I[Search and vector indexes]

Raw documents and extraction records belong in durable artifact/storage systems. A vector index is a derived search structure, not the sole copy of the extracted handbook. See chunking strategies.

Add detail where the baseline fails

Failure Improvement Cost/benefit
One bad page fails the full upload Page-level status and selective retries More job bookkeeping; much less repeated work
Retried pages create duplicate chunks Keys include document revision, page and pipeline version Requires careful replacement/index cleanup
Hundreds of uploads exhaust provider limits Global and per-account concurrency/rate budgets Queueing adds delay but protects service stability
A table crosses a page boundary Boundary-aware assembly after page extraction More document-level processing
A model omits difficult text Compare region/page coverage; mark uncertainty Additional checks or review cost
Deleted documents remain searchable Propagate deletion/access revocation to every derived store Lifecycle tracking across artifacts and indexes
Source contains adversarial instructions Treat extracted text as data; keep tool permissions outside it Requires a clear trust boundary in later agents

A document can finish with explicitly reported partial failures. Do not display “complete” when unreadable pages were silently skipped.

Capacity and economics: calculate before promising speed

Suppose an illustrative extractor takes four seconds per page and allows 20 concurrent page requests. With equal page times and no other bottleneck:

500 pages / 20 concurrent requests = 25 waves
25 waves × 4 seconds               = 100 seconds of extraction

Upload, rendering, queueing, rate limits, retries and assembly add time. A ten-page request can have different cost/latency from ten one-page requests. Measure the selected provider's behavior and limits; “50 workers” alone does not imply completion in 20 seconds.

For a hypothetical hybrid pipeline over 1,000 pages:

base extraction: 1,000 × $0.002          = $2.00
vision fallback: 100 difficult pages × $0.03 = $3.00
total extraction charge                = $5.00

These are assumed unit costs, not current vendor quotes. Add storage, rendering, orchestration and human review. Compare against quality: the cheapest pipeline that drops table rows may create a more expensive downstream failure. For vision APIs, image size/detail, model and output length can affect billing; a universal “price per page” is misleading.

Evaluate transcription and structure separately

Character error rate (CER) is (substitutions + deletions + insertions) / reference characters under the chosen alignment and normalization. Word error rate (WER) applies the analogous calculation to words. State how whitespace, punctuation and Unicode are normalized.

Example: 10 substitutions, 5 deletions and 2 insertions against 1,000 reference characters gives 17 / 1,000 = 1.7% CER. CER can exceed 100% when insertions are large; it is not simply a bounded “accuracy percentage.” A low average CER can still conceal a dangerous decimal-point or identifier error.

Quality dimension Useful measure or review
Text CER/WER and exact match for critical identifiers/values
Reading order Correct region sequence on labeled layouts
Tables Cell text, row/column assignment, headers and span correctness
Evidence Correct source page/region and usable citation
Coverage Missing pages, blocks, table rows and captions
Downstream use Retrieval/answer correctness on document-grounded questions
Operations p95 job time, failures, retries, review rate and cost per accepted page

Use a held-out set containing clean PDFs, scans, handwriting, multiple languages, rotated text and complex tables. Tune review thresholds on that set. Textract's confidence guidance recommends considering both scores and use-case sensitivity. Check score calibration on your data; scores from different engines are not automatically comparable.

Interview questions and answer checks

  1. Why not send every page to a vision model? Native extraction may preserve more detail at lower cost; difficult regions can justify a selective vision path.
  2. Is traditional OCR deterministic and therefore correct? Repeatability, when achieved, does not establish correctness. OCR can substitute, omit or insert characters.
  3. Why can good OCR produce bad RAG answers? Reading order, table headers, chunk boundaries, units or source access may be wrong even when characters are right.
  4. How do you handle handwriting? Evaluate the actual language/style and image quality; use an appropriate engine and uncertainty/review policy rather than assuming human-level accuracy.
  5. What should be cached? Versioned extraction artifacts keyed by source and pipeline configuration, with the same access/deletion rules as the document.
  6. How do you recover after page 327 fails? Keep completed page results, retry the failed page within limits, and rerun affected assembly/index stages.
  7. When is Markdown insufficient? When merged cells, figures, exact coordinates, reading-order relations or evidence need richer structure.
  8. What is the closing design argument? Preserve source evidence, route by document characteristics, validate structure and completeness, and spend extra compute/review where measured errors justify it.

Final notes

Remember recognize → structure → verify → cite. OCR and vision are extraction tools, not authorities on what the source must have meant. Preserve the document, preserve uncertainty, and measure the errors that would matter to the application.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Framework changes: reproduce, diagnose, migrate
NEXT LESSONLLM infrastructure: size work, protect deadlines, recover failures →

Explore the diagram