LlamaIndex is a framework for connecting applications to external data, especially for retrieval-augmented generation (RAG). It provides document ingestion, indexing, retrieval, query engines, and agent integrations. Its Workflows library adds event-driven application control. A framework supplies useful components; the application still owns data permissions, correctness, recovery, and operating cost.
For a Learnastra interview, explain the document lifecycle before naming the framework. An excellent retriever cannot repair a table whose parser assigned an amount to the wrong row. A well-parsed document still must not reach an unauthorized reader.
The core concepts
| Concept | Plain definition | Example in an interview-preparation library |
|---|---|---|
| Document | A source item with content and metadata | One version of a chapter |
| Node | A unit derived from a document, often a text chunk | A section on partition failures, linked to its source |
| Index | A structure that supports finding relevant data | Vectors plus references to source nodes |
| Retriever | A component that selects candidate nodes for a query | Return passages explaining linearizability |
| Query engine | A higher-level interface that retrieves and produces a response | Answer with citations from the selected passages |
| Agent | A model-directed loop that can choose tools | Choose search, then fetch a cited section |
| Workflow | Steps connected through typed events and execution rules | Parse, validate, index, and publish a document revision |
A node need not be a fixed-size text fragment. Its useful fields include stable identity, source revision, section/page location, and access scope. Preserve the relationship to the source when transforming content. See the official document and node concepts.
VectorStoreIndex and PropertyGraphIndex remain framework-core abstractions. The presence of separate integrations and a standalone Workflows package does not mean indexing has disappeared from the core library. Check the vector index documentation for the installed version.
Interview exercise: a versioned document assistant
Functional requirements
- Ingest approved documents and retain source identifiers and revisions.
- Answer questions using only documents the signed-in reader can access.
- Return citations that open the precise source section.
- Replace or remove outdated content without leaving searchable stale fragments.
- Report indexing failures and abstain when evidence is insufficient.
Non-functional requirements
- Enforce access control during retrieval and again when opening a source.
- Make an ingestion retry safe and expose the current published revision.
- Set separate latency targets for interactive questions and background ingestion.
- Bound parsing, embedding, and generation expenditure.
- Measure retrieval quality and answer grounding on held-out questions.
These are application requirements, not guarantees provided by choosing LlamaIndex. Start with a small query engine and explicit metadata filtering. Add workflows when ingestion has branches, asynchronous work, or recovery requirements that warrant them.
Read diagram source
flowchart TD
A[Approved source revision] --> B[Parse and preserve layout]
B --> C{Content and metadata valid?}
C -->|No| D[Record failure for review]
C -->|Yes| E[Create nodes and embeddings]
E --> F[Stage index revision]
F --> G[Publish validated revision]
Q[Authenticated question] --> R[Derive access scope]
R --> S[Retrieve from published index]
G --> S
S --> T{Sufficient relevant evidence?}
T -->|No| U[Explain evidence gap]
T -->|Yes| V[Generate and validate citations]
Start simple, then repair the failure modes
Baseline: load documents, split them into nodes, embed the nodes, and retrieve a small candidate set. A query engine supplies that evidence to a model. Use explicit storage configuration for production; an in-memory demonstration does not establish persistence.
| Failure observed | Improvement | Benefit | Cost or limitation |
|---|---|---|---|
| Heading separated from the paragraph it qualifies | Preserve document structure or add parent context | Better interpretation | Larger context and more preprocessing |
| Exact product identifiers are missed | Combine lexical and vector retrieval | Better identifier recall | Candidate merging and tuning |
| Retrieved text is obsolete | Filter by published revision and reconcile deletions | Fresher answers | Version lifecycle and reconciliation jobs |
| Model confuses columns in a table | Improve parsing and test row/column fidelity | More reliable evidence | Parsing expense and manual review |
| Access changes after indexing | Apply current access policy at query time | Avoid relying on stale ACL metadata alone | Authorization lookup and cache invalidation |
| Every question queries every corpus | Add a tested query router | Lower unnecessary work | Routing errors can hide relevant evidence |
Choose semantic splitting because evaluations show it improves your corpus. It is not inherently optimal, and not every semantic splitter uses an LLM: embedding-based similarity is another approach. Compare against a structure-aware baseline. See chunking strategies.
Make ingestion an explicit lifecycle
LlamaIndex's ingestion pipeline can use a document store to track document identities and content hashes. With an attached vector store, changed documents can be reprocessed and upserted. The exact strategy depends on configuration; this is not automatic correctness for every connector. See the ingestion pipeline.
For the library example:
- Assign a stable logical document ID and an immutable revision ID.
- Record the parser, splitter, and embedding configurations used for that revision.
- Parse and validate before publishing the new index view.
- Retry failed stages using stable job and node identities.
- Publish a revision only after its required artifacts are ready.
- Remove obsolete nodes, reconcile source deletions, and invalidate related caches.
A source hash detects changed bytes. An unchanged source may still need processing after a parser bug fix or embedding-model change. Include those dependencies in the decision to rebuild. A database upsert does not itself make several storage systems update atomically.
Capacity example: assume 2,000 documents, 20 pages per document, and 3 nodes per page. That produces 120,000 nodes. At 768 float32 dimensions, the raw vectors occupy 120,000 × 768 × 4 = 368,640,000 bytes, about 369 MB. This excludes text, metadata, index overhead, replication, and backups. Estimate parsing cost by the provider's actual billing unit; do not confuse pages with embedding tokens.
When a property graph helps
A property graph represents entities and relationships, with properties on either. A query such as “Which services depend on a library with this vulnerability?” may benefit from traversing verified dependency edges. LlamaIndex supports property-graph indexing and retrieval.
Graph extraction can introduce wrong entities or relationships. Each inferred edge needs provenance, and extraction changes can require rebuilding graph data. Graphs add operational and evaluation work.
| Query | First approach to consider | Reason |
|---|---|---|
| Find documents by author and date | Metadata or relational filters | Explicit fields already answer it |
| Explain a concept in different wording | Vector or hybrid retrieval | Semantic matching is useful |
| Follow several dependency relationships | Graph traversal plus source retrieval | Relationships define the query |
| Summarize an entire large collection | A separately evaluated aggregation strategy | Top-k passages may omit most of the corpus |
A property graph is an option with a specific purpose, not an automatic upgrade over vector retrieval.
Workflows: events, state, and concurrency
Workflows routes typed events to steps that accept those event types. The standalone package uses imports from workflows; compatibility surfaces also exist within LlamaIndex. Follow one supported API version consistently. See the Workflows documentation.
For document ingestion, events might represent ParsedDocument, ValidatedDocument, IndexWriteResult, and RejectedDocument. The application must decide which transitions are allowed and whether all required work succeeded before publication.
| Concern | Design decision |
|---|---|
| Shared state | Keep small, typed state; use the context store and its supported atomic edit mechanism |
| Fan-out | Bound workers and queue size; enforce provider quotas across workers |
| Join | Identify the expected jobs and reject incomplete results when completeness is required |
| Ordering | Carry document/node IDs; completion order is not source order |
| Streaming | Send progress events; label partial results as incomplete |
| Human review | Persist the pending revision and resume only with an authorized decision |
Current documentation describes finite batch fan-out/join patterns and worker limits. Async execution helps overlap waiting for I/O; it does not make CPU-heavy parsing automatically parallel or establish a cluster-wide limit. See concurrent execution and state management.
Saving context is only part of recovery
A serialized context must be written to durable storage at useful boundaries and restored correctly. An occasional manual snapshot can lose work performed afterward. External operations also need stable operation IDs or reconciliation: a restored workflow cannot infer whether an index write completed just before a crash. Review the durable workflows guide and durable execution concepts.
Prefer supported safe serialization for application state. Never unpickle an untrusted checkpoint. Saving state, resuming computation, and preventing duplicate external effects are separate responsibilities.
Managed parsing and framework composition
LlamaIndex's managed document products provide parsing, extraction, and indexing capabilities. Product names and packages evolve; evaluate the current document platform rather than relying on a fixed list of internal model providers.
Compare managed parsing against local parsing on representative files: multi-column text, merged table cells, scanned pages, handwriting where relevant, and difficult layouts. Record accuracy, turnaround time, retries, data handling, and cost per successfully processed document. A confidence field is useful only if its calibration is adequate for your acceptance decision.
A query engine can also be exposed as a tool to a different agent runtime. Give that tool a bounded contract: authorized corpus scope, query, evidence budget, citations, freshness, and explicit failure results. The calling agent should not receive arbitrary database credentials.
| Selection question | Practical answer |
|---|---|
| Does the application mainly need ingestion and retrieval components? | Evaluate LlamaIndex on that corpus |
| Are graph state transitions and checkpoint inspection already central to the application? | Compare the existing LangGraph setup against Workflows with a recovery exercise |
| Must two frameworks be combined? | Only when the extra capability exceeds adapter, tracing, and versioning costs |
| Can a Python design move unchanged to TypeScript? | Verify feature and integration availability; shared branding does not guarantee parity |
Interview questions and answer notes
- Why can an unchanged document need reindexing? A transformation or embedding version changed. Source bytes are only one dependency.
- Does a property graph solve “documents by author last month” better? Not necessarily. Indexed metadata may provide a simpler, exact answer.
- An async workflow exceeds a provider quota. Why? Async does not establish a global concurrency or request-rate budget. Enforce both where required.
- A checkpoint exists, but duplicate writes appear after recovery. What is missing? An idempotent write protocol or reconciliation between workflow state and the external system.
- Would you deploy LlamaIndex plus LangGraph by default? No. First identify the capability each contributes and test the integration's failure boundaries.
- How do you verify a PDF parsing improvement? Compare grounded downstream answers and document-level structural accuracy on a fixed sample, along with cost and latency.
- Why is a citation insufficient evidence of correctness? It may refer to an irrelevant, obsolete, unauthorized, or misparsed source. Validate support and source identity.
Final notes
Remember source → nodes → index → retrieval → supported answer. Workflows coordinates the surrounding process. In an interview, close with the revision lifecycle, authorization boundary, and recovery test that make the design dependable. Framework choice follows those requirements.