Day 2: Hybrid Search and Reranking for Production RAG (Dense + BM25 + Cross-Encoders)

Production RAG Masterclass · Day 2 of 7

If your RAG system “almost” finds the right answer when users paste an error code, a part number, or an acronym — the problem is rarely the LLM. It is almost always retrieval that smoothed away the exact tokens that mattered.

Hybrid search combining keyword BM25 retrieval with dense vector search and reranking

Figure 1. Sparse and dense retrieval merged before a final rerank pass.

Thesis Vector distance is great at meaning and weak at literals. Hybrid search closes that gap; a cross-encoder decides what actually belongs in the prompt.

Where dense-only retrieval quietly fails

Embedding models shine on paraphrases. Production users do not. They paste strings the model was never trained to treat as sacred:

  • ERR_TIMEOUT_5042
  • PN-88421-B
  • Domain shorthand: SOX, MTTR, OIDC, SKU families

Cosine similarity softens those tokens into “something nearby.” BM25 keeps them exact. You need both lists, then a principled way to merge them.

Stage 1: run sparse and dense in parallel

Dense path — embed the query, nearest-neighbor over chunk vectors.

Sparse path — BM25 (or a lexical field) over the same corpus.

Keep a combined candidate pool of roughly 20–40 chunks before Stage 2. That range usually balances recall against rerank latency.

Retrieval pipeline showing BM25, dense vectors, Reciprocal Rank Fusion, and cross-encoder reranking

Figure 2. Fuse ranks first. Rescore second. Generate last.

Reciprocal Rank Fusion (RRF)

Do not average BM25 scores with cosine scores. The scales are incompatible. Fuse on rank instead:

RRF(d) = Σ  1 / (k + rank_i(d))

# k ≈ 60 is the common default
# rank_i(d) = 1-indexed position in retriever i

A document near the top of either list rises. A document near the top of both rises faster. That is the hybrid win in one formula.

Minimal Python sketch

from collections import defaultdict

def rrf(rank_lists, k=60):
    scores = defaultdict(float)
    for ranks in rank_lists:  # each: [doc_id, ...] best → worst
        for r, doc_id in enumerate(ranks, start=1):
            scores[doc_id] += 1.0 / (k + r)
    return sorted(scores, key=scores.get, reverse=True)

fused = rrf([bm25_hits, dense_hits])[:40]

Stage 2: cross-encoder reranking

Bi-encoders embed the query and the document separately — fast enough for an index, shallow on token interaction. Cross-encoders such as bge-reranker-large score the pair jointly.

Feed the fused Top-N into the reranker. Pass only the final Top-K (often 5–10) into the LLM context.

pairs = [(query, chunk.text) for chunk in fused_candidates]
scores = cross_encoder.predict(pairs)
top_k = [c for _, c in sorted(zip(scores, fused_candidates), reverse=True)[:8]]

Latency knobs that actually matter

  • Candidate pool: start at 20–40; measure p95 before growing it.
  • Reranker size: base when latency is tight; large when quality wins.
  • Sparse weight: lean BM25 for ID-heavy traffic; lean dense for conversational asks.
  • Cache: memoize rerank results for repeated internal FAQs.

What good looks like after Day 1

Once chunking is honest (Day 1), Day 2 usually recovers the failures where users paste ticket IDs or product codes and get paragraphs that feel semantically close but operationally wrong.

Key takeaways

  1. Pair dense vectors with BM25 when exact tokens matter.
  2. Merge with RRF — not raw score averaging.
  3. Rerank a small pool with a cross-encoder before generation.
  4. Cap candidates and watch p95 latency as you tune.

Series: Production RAG Masterclass
Previous: Day 1 — Ingestion & Chunking
Next: Day 3 — Query Optimization (HyDE & Expansion) 

Comments

Popular posts from this blog

LangChain 2.0 Tutorial: Build an AI Agent with Tools (2025 Edition)

Best AI Tools for Business Owners in 2025: Your Secret Weapon for Super Productivity & More Free Time!

FastAPI for AI Developers: Build Lightning-Fast APIs in 2025