Day 5: Agentic RAG: Query Routing, Self-Reflection, and Corrective RAG (CRAG)

Production RAG Masterclass · Day 5 of 7

A linear pipeline assumes the first retrieval was good enough. Production questions often need a second look, a different tool, or an honest “we do not have evidence.”

Agentic RAG loop with routing, reflection, and correction

Figure 1. Plan, retrieve, critique, then answer — or try again.

Thesis Static, single-shot retrieval fails when a system must verify sources or gather multi-step evidence. Move from a pipeline to a bounded loop.

From pipeline to loop

Classic RAG: query → retrieve → generate → done.

Agentic RAG: plan → route → retrieve → critique → (retry | answer) with a hard stop on iterations.

Retrieve reflect correct agent loop diagram for RAG

Figure 2. Cap loops so latency stays predictable.

Query routing

Not every question belongs in the vector DB:

  • Vector index — policies, runbooks, wiki.
  • SQL / warehouse — metrics, counts, “how many…”.
  • External API — live status, tickets, prices.
  • Graph — multi-hop entity questions (Day 4).
route = llm_json(f"""Choose tools for this question.
Options: vector, sql, api, graph
Question: {q}
Return {{"tools": [...], "reason": "..."}}""")

Self-correction and reflection

After a draft answer, run a critic pass:

  1. Does every claim cite a retrieved chunk?
  2. Are any statements unsupported?
  3. Did we answer the actual question?
critique = llm(f"""Given SOURCES and ANSWER, list unsupported claims.
SOURCES:
{ctx}
ANSWER:
{draft}""")
if critique.has_gaps:
    q2 = rewrite_query(q, critique)
    ctx += retrieve(q2)

Corrective RAG (CRAG)

Score retrieval confidence. If low:

  • Rewrite the query and search again, or
  • Fall back to an approved external search, or
  • Refuse with a clear “insufficient evidence” response.

CRAG is the production pattern for “do not invent when the corpus is empty.”

Guardrails for agent loops

  • Max 2–3 retrieval iterations.
  • Budget tokens and tool calls per request.
  • Log every route and critique for evaluation (Day 7).
  • Never let the agent invent tool results.

Key takeaways

  1. Replace one-shot RAG with plan → retrieve → reflect loops.
  2. Route to vector, SQL, API, or graph based on the ask.
  3. Critique answers for attribution before returning them.
  4. CRAG: rewrite or fall back when retrieval confidence is low.

Series: Production RAG Masterclass
Previous: Day 4 — GraphRAG
Next: Day 6 — Enterprise Security

Comments

Popular posts from this blog

LangChain 2.0 Tutorial: Build an AI Agent with Tools (2025 Edition)

Best AI Tools for Business Owners in 2025: Your Secret Weapon for Super Productivity & More Free Time!

Unlocking the Future: 10 Key Insights into Web3 Technologies