Posts

Day 4: GraphRAG Explained: Combining Knowledge Graphs with Vector Search

Image
Production RAG Masterclass · Day 4 of 7 Some questions are not “find the nearest paragraph.” They are “walk a path across three documents and explain the join.” Flat vector search was never designed for that. Figure 1. Vectors find seeds. Graphs find relationships. Thesis Cosine similarity cannot infer multi-hop entity relationships across separate documents. Pair vectors with an entity–relation graph. Where flat vector RAG breaks Ask: “Which vendors touch the same data pipeline as Team Atlas, and who owns the SLA?” The facts live in an org chart, a vendor contract, and a runbook. No single chunk holds the full path. Similarity returns locally relevant paragraphs — not the chain. Build the graph at ingest Extract triples while you chunk: (Entity) -[RELATION]-> (Entity) (Team Atlas)-[OWNS]->(Billing Pipeline) (Billing Pipeline)-[DEPENDS_ON]->(Vendor Stripe) (Vendor Stripe)-[HAS_SLA]->(99.95%) Store nodes and edges in Neo4j (or similar). Keep chunk t...

Day 3: Query Optimization for RAG: HyDE, Multi-Query Expansion, and Sub-Query Decomposition

Image
Production RAG Masterclass · Day 3 of 7 Most RAG post-mortems blame the index. In practice, the query that hit the index was a vague chat message that never resembled the documents you stored. Figure 1. Fix the query shape before you retune the vector store. Thesis Raw user queries are short, vague, and poorly formatted for semantic matching. Optimize the query — then retrieve. The query–document mismatch User: “why is billing failing?” Doc: “Invoice reconciliation jobs abort when Stripe webhook signature validation returns HTTP 400…” Those strings barely share tokens. Embeddings help. A deliberate bridge helps more. Strip conversational noise first Before expansion, normalize the prompt: Drop politeness filler (“please”, “can you”, “thanks”). Drop chat history that is not about the current ask. Keep entities, IDs, product names, and constraints. def clean_query(q: str) -> str: return llm(f"Rewrite as a concise search query:\n{q}") ...

Day 2: Hybrid Search and Reranking for Production RAG (Dense + BM25 + Cross-Encoders)

Image
Production RAG Masterclass · Day 2 of 7 If your RAG system “almost” finds the right answer when users paste an error code, a part number, or an acronym — the problem is rarely the LLM. It is almost always retrieval that smoothed away the exact tokens that mattered. Figure 1. Sparse and dense retrieval merged before a final rerank pass. Thesis Vector distance is great at meaning and weak at literals. Hybrid search closes that gap; a cross-encoder decides what actually belongs in the prompt. Where dense-only retrieval quietly fails Embedding models shine on paraphrases. Production users do not. They paste strings the model was never trained to treat as sacred: ERR_TIMEOUT_5042 PN-88421-B Domain shorthand: SOX, MTTR, OIDC, SKU families Cosine similarity softens those tokens into “something nearby.” BM25 keeps them exact. You need both lists, then a principled way to merge them. Stage 1: run sparse and dense in parallel Dense path — embed the query, ne...

Day 1: Stop Letting Naive Chunking Ruin Your RAG Pipelines

Image
Part 1 of 7 — RAG Systems in Practice Stop Letting Naive Chunking Ruin Your RAG Pipelines When a RAG system answers a query with confidence and gets it totally wrong, the initial reaction is usually to blame the model's temperature or swap vector databases. Most of the time, the real bug happened hours earlier when you ingested the raw PDF. What actually breaks when you use fixed character splitting? Standard tutorials tell you to take your document and run a RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50) . It works fine for demo apps, but here is what happens on enterprise data: Tables turn into nonsense: A 10-column financial or spec table gets chopped right in half. Row entries lose their column headers, turning structured data into random numbers. Headers detach from paragraphs: The section heading ### Rate Limits lands in Chunk A, while the actual rate l...

Why Small Language Models Are Becoming the Brain of AI Agents in 2026

Image
Why Small Language Models Are Becoming the Brain of AI Agents in 2026 AI agents are evolving rapidly in 2026. From automation assistants to coding copilots and workflow bots, modern AI systems are no longer limited to simple chat interfaces. But something interesting is happening behind the scenes — developers are now shifting toward Small Language Models (SLMs) for tool calling and intelligent automation workflows. What Are AI Agents? An AI agent is a system that can: understand tasks make decisions call tools or APIs remember context continue workflows automatically Unlike traditional chatbots , agents can actually perform actions instead of only generating responses. User Request → AI Agent → Tool/API → Result → Next Action Why Small Language Models Are Trending For years, the AI industry focused mainly on large models with massive infrastructure requirements. But in real-world automation systems, developers realized something imp...

10 Common Python Mistakes Every Developer Makes (And How to Avoid Them in 2026)

Image
10 Common Python Mistakes Every Developer Makes (And How to Avoid Them in 2026) Even experienced developers make small mistakes that lead to bugs, slow performance, or unreadable code. Below are 10 common Python mistakes—and the simple, correct patterns you should use instead. 1. Using Mutable Default Arguments Defining a function with a mutable default (like a list or dict) can cause the same object to persist across calls. def add_item(item, items=[]): items.append(item) return items # BAD: same list reused across calls Use None as the default and create a new object inside the function. def add_item(item, items=None): if items is None: items = [] items.append(item) return items # Good: fresh list each call 2. Misunderstanding Shallow vs Deep Copy Assigning one list to another does not copy it; both names reference the same object. list2 = list1 # Not a copy — both reference the same list Use copy() for a shallow copy o...

NeuroFlow Python Scripts — Using Lightweight Neural Models for Local Automation

Image
NeuroFlow Python Scripts — Using Lightweight Neural Models for Local Automation AI automation usually depends on cloud services like OpenAI , AWS , or Google APIs . But in 2025, a new approach is rising — NeuroFlow Python Scripts , where small, lightweight neural models run locally on your system to automate tasks, predict actions, and make intelligent decisions without the cloud. This is a brand-new concept: local AI-driven automation that works offline, consumes low memory, and learns your patterns over time. What Are NeuroFlow Python Scripts? NeuroFlow Scripts are Python automation scripts enhanced with: tiny neural models (under 5–20MB) local inference without cloud APIs pattern recognition from your daily tasks adaptive actions based on usage history context-aware decisions Think of it as “ mini AI ” inside your automation scripts. Why NeuroFlow-Based Automation? No cloud dependency No API cost Runs offline Faster on ...