Agentic RAG uses model-driven decisions to choose retrieval actions as part of an answering workflow. The model may select a source, formulate a follow-up query or decide that more evidence is needed. Application code still controls available tools, permissions, budgets and completion criteria.
The distinction from a fixed RAG pipeline is who chooses the next step. A fixed pipeline can include query rewriting, multiple searches, validation and recovery. An agentic pipeline does not automatically have higher quality or stronger self-correction.
Use it for a variable evidence path
Consider a hypothetical internal question: “Which team owns the service that calculates tax for Checkout?” Available sources contain:
- A service registry linking Checkout to TaxCalc.
- A current ownership directory linking TaxCalc to the Platform team.
The second lookup depends on the identifier found in the first. The system must verify both relationships and their current versions. If the registry says the tax dependency was removed, the premise changes; the agent should correct the premise instead of inventing the expected second hop.
For this particular task, a direct join or fixed two-step workflow might be simpler. Agentic retrieval is worth evaluating when source choice and follow-up needs vary across a broader workload.
Define the requirements
Functional requirements:
- Resolve the user's question and any missing constraints.
- Choose among permitted retrieval tools and follow evidence-dependent steps.
- Preserve source IDs and applicability for each supported claim.
- Answer, clarify or report insufficient evidence with an explicit outcome.
Non-functional requirements:
- Limits on tool calls, tokens, elapsed time and concurrent work.
- No exposure of unauthorized content or private queries to public search.
- Traceable decisions and recoverable state across retries or worker failures.
- Measured end-to-end quality, latency, cost and no-progress behavior.
Build a bounded loop
Read diagram source
flowchart TD
Q[Question and authenticated scope] --> P[Choose permitted retrieval step]
P --> B{Budget and policy allow it?}
B -->|Yes| T[Execute retrieval]
T --> V[Validate results and update evidence state]
V --> E{Enough supported evidence?}
E -->|Yes| A[Answer with sources]
E -->|No and new useful search exists| P
E -->|Missing user input| C[Ask targeted clarification]
B -->|No| S[Return explicit incomplete outcome]
E -->|No progress| S
The stop conditions are enforced by code. A model-generated “continue” decision cannot reset the budget. Exhausting the budget is not evidence that the remaining uncertainty has disappeared.
Keep application-visible state such as:
{
"question_id": "q-17",
"resolved_service": "TaxCalc",
"evidence_ids": ["registry-v42:checkout-tax"],
"unresolved": ["current owning team"],
"completed_queries": ["Checkout dependency registry"],
"retrieval_calls_used": 1,
"status": "gathering_evidence"
}
This is task state, not hidden model reasoning. Store durable operation state separately when recovery requires it. See context engineering.
Related research approaches are not interchangeable
| Approach | Distinguishing mechanism | Production question |
|---|---|---|
| Self-RAG | Trains a model to use retrieval and reflection tokens for generation/critique | Are the trained model and inference policy available and suitable? |
| Corrective RAG (CRAG) | Evaluates retrieval quality and routes through corrective processing, including external search in the studied design | Is the evaluator reliable, and is the fallback source permitted? |
| Adaptive-RAG | Selects retrieval complexity according to the question | How often does the router choose an insufficient path? |
| ReAct-style retrieval | Alternates model-selected actions and observations | What bounds prevent loops, scope expansion and unsupported conclusions? |
Self-RAG is not merely prompting any model to print “Relevant” or “Supported.” Those labels do not provide a correctness certificate. CRAG's proposed external-search path must be adapted to the application's privacy and source requirements. A private document question must not automatically be sent to a public search engine after a weak retrieval score. Self-RAG, CRAG, Adaptive-RAG, ReAct.
Plan, revise and parallelize carefully
Orchestration libraries such as LangGraph and LlamaIndex can help implement stateful workflows and tool execution. Choose based on checkpointing, recovery, integration and operational needs. A library does not supply the application's evidence standard or authorization policy. LangGraph overview, LlamaIndex agents.
Separate independent subqueries from dependent hops. Two policy regions may be searched concurrently; the owner lookup in the service example cannot use the correct service ID until the dependency lookup resolves it.
Speculative retrieval can start likely future searches early, but spends resources on guesses. It must not bypass access checks or allow unverified intermediate assumptions to become facts. Measure wasted calls and stale results alongside latency savings.
Replanning should respond to evidence: a missing source, a corrected identifier or an unresolved conflict. Rephrasing the same failing query indefinitely is not useful progress. Track repeated queries, repeated result IDs and changes in evidence coverage.
Compare cost with a simpler baseline
An illustrative three-round sequence might spend 400 ms choosing each step and 150 ms retrieving, followed by 800 ms for the answer: 3 × (400 + 150) + 800 = 2,450 ms, excluding queues and network overhead. These are assumed durations, not measured model performance.
Track total tokens across every planner, evaluator, retrieval-expansion and answer call. A final answer using 500 tokens does not mean the task cost only 500 tokens. Compare cost per adequately supported answer and severe-error rate with a fixed workflow.
| Failure | Repair | Tradeoff |
|---|---|---|
| Repeated searches without new evidence | No-progress detection and stop policy | Some difficult searches stop early |
| Wrong intermediate entity | Validate identifier and source before next hop | Extra lookup or clarification |
| Evaluator approves weak evidence | Independent evidence checks and calibrated evaluation | More validation work |
| Public fallback leaks private intent | Source policy and data-use restrictions | Fewer fallback options |
| Tool text redirects the agent | Preserve provenance and enforce capabilities | Model-side detection remains fallible |
| Long trajectory loses constraints | Structured state and careful compaction | State maintenance and verification |
Interview practice
Q1: When is agentic retrieval preferable to fixed RAG?
When questions require variable source selection or evidence-dependent follow-ups that the fixed flow handles poorly. I would demonstrate the failure class and compare the complete agentic path against a simpler workflow, including cost and latency.
Q2: Does a fixed workflow have no self-correction?
No. It can contain validation, retries and alternate retrieval branches. Agentic describes model discretion over actions; reliability depends on the whole design and its acceptance checks.
Q3: What should happen after the maximum number of searches?
Return a supported answer if the evidence is sufficient, otherwise report what remains unresolved or ask for required input. Do not force an unsupported final answer merely because a turn limit was reached.
Q4: How do you evaluate a multi-hop answer?
Check each intermediate entity and relationship, source applicability and the final synthesis. Include questions with false premises or absent links. A fluent final sentence can conceal an incorrect first hop.
Q5: Does choosing an orchestration framework make the agent deterministic?
It can constrain allowed transitions and provide durable state, but model decisions may still vary. Version the workflow, model and inputs, and evaluate trajectory and outcome variability. Framework selection is an implementation decision after the behavior is defined.
Q6: How do you reduce latency safely?
Route suitable simple tasks to a measured fast path, parallelize independent retrievals, reuse permitted fresh results and limit unnecessary critique calls. Evaluate router mistakes and speculative waste; removing checks indiscriminately can produce faster unsupported answers.
Final notes
Recall card: Evidence-dependent steps, code-enforced bounds, supported outcomes. The agent's ability to choose another search is useful only if that search improves the task.
Related: GraphRAG, advanced retrieval patterns, prompt injection.