Posts

Showing posts with the label Reranking

Day 2: Hybrid Search and Reranking for Production RAG (Dense + BM25 + Cross-Encoders)

Image
Production RAG Masterclass · Day 2 of 7 If your RAG system “almost” finds the right answer when users paste an error code, a part number, or an acronym — the problem is rarely the LLM. It is almost always retrieval that smoothed away the exact tokens that mattered. Figure 1. Sparse and dense retrieval merged before a final rerank pass. Thesis Vector distance is great at meaning and weak at literals. Hybrid search closes that gap; a cross-encoder decides what actually belongs in the prompt. Where dense-only retrieval quietly fails Embedding models shine on paraphrases. Production users do not. They paste strings the model was never trained to treat as sacred: ERR_TIMEOUT_5042 PN-88421-B Domain shorthand: SOX, MTTR, OIDC, SKU families Cosine similarity softens those tokens into “something nearby.” BM25 keeps them exact. You need both lists, then a principled way to merge them. Stage 1: run sparse and dense in parallel Dense path — embed the query, ...