Day 2: Hybrid Search and Reranking for Production RAG (Dense + BM25 + Cross-Encoders)
Production RAG Masterclass · Day 2 of 7
If your RAG system “almost” finds the right answer when users paste an error code, a part number, or an acronym — the problem is rarely the LLM. It is almost always retrieval that smoothed away the exact tokens that mattered.
Figure 1. Sparse and dense retrieval merged before a final rerank pass.
Thesis Vector distance is great at meaning and weak at literals. Hybrid search closes that gap; a cross-encoder decides what actually belongs in the prompt.
Where dense-only retrieval quietly fails
Embedding models shine on paraphrases. Production users do not. They paste strings the model was never trained to treat as sacred:
ERR_TIMEOUT_5042PN-88421-B- Domain shorthand: SOX, MTTR, OIDC, SKU families
Cosine similarity softens those tokens into “something nearby.” BM25 keeps them exact. You need both lists, then a principled way to merge them.
Stage 1: run sparse and dense in parallel
Dense path — embed the query, nearest-neighbor over chunk vectors.
Sparse path — BM25 (or a lexical field) over the same corpus.
Keep a combined candidate pool of roughly 20–40 chunks before Stage 2. That range usually balances recall against rerank latency.
Figure 2. Fuse ranks first. Rescore second. Generate last.
Reciprocal Rank Fusion (RRF)
Do not average BM25 scores with cosine scores. The scales are incompatible. Fuse on rank instead:
RRF(d) = Σ 1 / (k + rank_i(d)) # k ≈ 60 is the common default # rank_i(d) = 1-indexed position in retriever i
A document near the top of either list rises. A document near the top of both rises faster. That is the hybrid win in one formula.
Minimal Python sketch
from collections import defaultdict
def rrf(rank_lists, k=60):
scores = defaultdict(float)
for ranks in rank_lists: # each: [doc_id, ...] best → worst
for r, doc_id in enumerate(ranks, start=1):
scores[doc_id] += 1.0 / (k + r)
return sorted(scores, key=scores.get, reverse=True)
fused = rrf([bm25_hits, dense_hits])[:40]
Stage 2: cross-encoder reranking
Bi-encoders embed the query and the document separately — fast enough for an index, shallow on token interaction. Cross-encoders such as bge-reranker-large score the pair jointly.
Feed the fused Top-N into the reranker. Pass only the final Top-K (often 5–10) into the LLM context.
pairs = [(query, chunk.text) for chunk in fused_candidates] scores = cross_encoder.predict(pairs) top_k = [c for _, c in sorted(zip(scores, fused_candidates), reverse=True)[:8]]
Latency knobs that actually matter
- Candidate pool: start at 20–40; measure p95 before growing it.
- Reranker size: base when latency is tight; large when quality wins.
- Sparse weight: lean BM25 for ID-heavy traffic; lean dense for conversational asks.
- Cache: memoize rerank results for repeated internal FAQs.
What good looks like after Day 1
Once chunking is honest (Day 1), Day 2 usually recovers the failures where users paste ticket IDs or product codes and get paragraphs that feel semantically close but operationally wrong.
Key takeaways
- Pair dense vectors with BM25 when exact tokens matter.
- Merge with RRF — not raw score averaging.
- Rerank a small pool with a cross-encoder before generation.
- Cap candidates and watch p95 latency as you tune.
Series: Production RAG Masterclass
Previous: Day 1 — Ingestion & Chunking
Next: Day 3 — Query Optimization (HyDE & Expansion)
Comments
Post a Comment