Posts

Day 5: Agentic RAG: Query Routing, Self-Reflection, and Corrective RAG (CRAG)

Image
Production RAG Masterclass · Day 5 of 7 A linear pipeline assumes the first retrieval was good enough. Production questions often need a second look, a different tool, or an honest “we do not have evidence.” Figure 1. Plan, retrieve, critique, then answer — or try again. Thesis Static, single-shot retrieval fails when a system must verify sources or gather multi-step evidence. Move from a pipeline to a bounded loop. From pipeline to loop Classic RAG: query → retrieve → generate → done . Agentic RAG: plan → route → retrieve → critique → (retry | answer) with a hard stop on iterations. Figure 2. Cap loops so latency stays predictable. Query routing Not every question belongs in the vector DB: Vector index — policies, runbooks, wiki. SQL / warehouse — metrics, counts, “how many…”. External API — live status, tickets, prices. Graph — multi-hop entity questions (Day 4). route = llm_json(f"""Choose tools for this question. Options: ...

Day 4: GraphRAG Explained: Combining Knowledge Graphs with Vector Search

Image
Production RAG Masterclass · Day 4 of 7 Some questions are not “find the nearest paragraph.” They are “walk a path across three documents and explain the join.” Flat vector search was never designed for that. Figure 1. Vectors find seeds. Graphs find relationships. Thesis Cosine similarity cannot infer multi-hop entity relationships across separate documents. Pair vectors with an entity–relation graph. Where flat vector RAG breaks Ask: “Which vendors touch the same data pipeline as Team Atlas, and who owns the SLA?” The facts live in an org chart, a vendor contract, and a runbook. No single chunk holds the full path. Similarity returns locally relevant paragraphs — not the chain. Build the graph at ingest Extract triples while you chunk: (Entity) -[RELATION]-> (Entity) (Team Atlas)-[OWNS]->(Billing Pipeline) (Billing Pipeline)-[DEPENDS_ON]->(Vendor Stripe) (Vendor Stripe)-[HAS_SLA]->(99.95%) Store nodes and edges in Neo4j (or similar). Keep chunk t...

Day 3: Query Optimization for RAG: HyDE, Multi-Query Expansion, and Sub-Query Decomposition

Image
Production RAG Masterclass · Day 3 of 7 Most RAG post-mortems blame the index. In practice, the query that hit the index was a vague chat message that never resembled the documents you stored. Figure 1. Fix the query shape before you retune the vector store. Thesis Raw user queries are short, vague, and poorly formatted for semantic matching. Optimize the query — then retrieve. The query–document mismatch User: “why is billing failing?” Doc: “Invoice reconciliation jobs abort when Stripe webhook signature validation returns HTTP 400…” Those strings barely share tokens. Embeddings help. A deliberate bridge helps more. Strip conversational noise first Before expansion, normalize the prompt: Drop politeness filler (“please”, “can you”, “thanks”). Drop chat history that is not about the current ask. Keep entities, IDs, product names, and constraints. def clean_query(q: str) -> str: return llm(f"Rewrite as a concise search query:\n{q}") ...

Day 2: Hybrid Search and Reranking for Production RAG (Dense + BM25 + Cross-Encoders)

Image
Production RAG Masterclass · Day 2 of 7 If your RAG system “almost” finds the right answer when users paste an error code, a part number, or an acronym — the problem is rarely the LLM. It is almost always retrieval that smoothed away the exact tokens that mattered. Figure 1. Sparse and dense retrieval merged before a final rerank pass. Thesis Vector distance is great at meaning and weak at literals. Hybrid search closes that gap; a cross-encoder decides what actually belongs in the prompt. Where dense-only retrieval quietly fails Embedding models shine on paraphrases. Production users do not. They paste strings the model was never trained to treat as sacred: ERR_TIMEOUT_5042 PN-88421-B Domain shorthand: SOX, MTTR, OIDC, SKU families Cosine similarity softens those tokens into “something nearby.” BM25 keeps them exact. You need both lists, then a principled way to merge them. Stage 1: run sparse and dense in parallel Dense path — embed the query, ne...

Day 1: Stop Letting Naive Chunking Ruin Your RAG Pipelines

Image
Part 1 of 7 — RAG Systems in Practice Stop Letting Naive Chunking Ruin Your RAG Pipelines When a RAG system answers a query with confidence and gets it totally wrong, the initial reaction is usually to blame the model's temperature or swap vector databases. Most of the time, the real bug happened hours earlier when you ingested the raw PDF. What actually breaks when you use fixed character splitting? Standard tutorials tell you to take your document and run a RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50) . It works fine for demo apps, but here is what happens on enterprise data: Tables turn into nonsense: A 10-column financial or spec table gets chopped right in half. Row entries lose their column headers, turning structured data into random numbers. Headers detach from paragraphs: The section heading ### Rate Limits lands in Chunk A, while the actual rate l...

Why Small Language Models Are Becoming the Brain of AI Agents in 2026

Image
Why Small Language Models Are Becoming the Brain of AI Agents in 2026 AI agents are evolving rapidly in 2026. From automation assistants to coding copilots and workflow bots, modern AI systems are no longer limited to simple chat interfaces. But something interesting is happening behind the scenes — developers are now shifting toward Small Language Models (SLMs) for tool calling and intelligent automation workflows. What Are AI Agents? An AI agent is a system that can: understand tasks make decisions call tools or APIs remember context continue workflows automatically Unlike traditional chatbots , agents can actually perform actions instead of only generating responses. User Request → AI Agent → Tool/API → Result → Next Action Why Small Language Models Are Trending For years, the AI industry focused mainly on large models with massive infrastructure requirements. But in real-world automation systems, developers realized something imp...

10 Common Python Mistakes Every Developer Makes (And How to Avoid Them in 2026)

Image
10 Common Python Mistakes Every Developer Makes (And How to Avoid Them in 2026) Even experienced developers make small mistakes that lead to bugs, slow performance, or unreadable code. Below are 10 common Python mistakes—and the simple, correct patterns you should use instead. 1. Using Mutable Default Arguments Defining a function with a mutable default (like a list or dict) can cause the same object to persist across calls. def add_item(item, items=[]): items.append(item) return items # BAD: same list reused across calls Use None as the default and create a new object inside the function. def add_item(item, items=None): if items is None: items = [] items.append(item) return items # Good: fresh list each call 2. Misunderstanding Shallow vs Deep Copy Assigning one list to another does not copy it; both names reference the same object. list2 = list1 # Not a copy — both reference the same list Use copy() for a shallow copy o...