Day 3: Query Optimization for RAG: HyDE, Multi-Query Expansion, and Sub-Query Decomposition
Production RAG Masterclass · Day 3 of 7
Most RAG post-mortems blame the index. In practice, the query that hit the index was a vague chat message that never resembled the documents you stored.
Figure 1. Fix the query shape before you retune the vector store.
Thesis Raw user queries are short, vague, and poorly formatted for semantic matching. Optimize the query — then retrieve.
The query–document mismatch
User: “why is billing failing?”
Doc: “Invoice reconciliation jobs abort when Stripe webhook signature validation returns HTTP 400…”
Those strings barely share tokens. Embeddings help. A deliberate bridge helps more.
Strip conversational noise first
Before expansion, normalize the prompt:
- Drop politeness filler (“please”, “can you”, “thanks”).
- Drop chat history that is not about the current ask.
- Keep entities, IDs, product names, and constraints.
def clean_query(q: str) -> str:
return llm(f"Rewrite as a concise search query:\n{q}")
HyDE — Hypothetical Document Embeddings
Ask the model to write a short answer as if it already knew, embed that paragraph, and search with it. You are aligning query space with document space.
Figure 2. HyDE searches against a synthetic answer, not the raw chat message.
hyde_prompt = f"""Write a short technical paragraph that would
answer this question in our internal docs tone:
{clean_q}"""
hypo = llm(hyde_prompt)
hits = vector_search(embed(hypo), k=20)
HyDE works well on conceptual questions in a consistent domain. It fails when the model invents confident jargon — always fuse with the original query embedding or BM25 from Day 2.
Multi-query expansion
Generate 3–5 paraphrases. Retrieve for each. Union or RRF the hits.
variants = llm_json(f"""Generate 4 diverse search queries for:
{clean_q}
Return JSON list of strings.""")
rank_lists = [dense_search(v, k=15) for v in [clean_q, *variants]]
fused = rrf(rank_lists)[:30]
Diversity matters: one keyword-heavy variant, one synonym-heavy, one procedural phrasing.
Sub-query decomposition
Multi-part questions need parallel retrieval, not one mushy embedding.
“Compare our SSO options and list the SLA for each IdP.”
- What SSO / IdP options exist?
- What is the SLA for each IdP?
Retrieve per sub-query, then write a structured answer that cites both evidence sets.
When to use which
- Noise filter: always.
- Multi-query: default for ambiguous short questions.
- HyDE: conceptual / how-it-works questions.
- Decomposition: compare / multi-hop / multi-constraint asks.
Key takeaways
- Clean conversational fluff before vector search.
- HyDE embeds a synthetic answer to close the query–doc gap.
- 3–5 query variants + RRF boost recall cheaply.
- Decompose multi-part questions into parallel retrievals.
Series: Production RAG Masterclass
Previous: Day 2 — Hybrid Search & Reranking
Next: Day 4 — GraphRAG
Comments
Post a Comment