RAG Rerankers
If you've built a retrieval-augmented generation (RAG) pipeline and found that answers are occasionally missing obvious context even though the right document is "in the knowledge base," the problem is usually the retrieval step, not the LLM.
Standard vector search relies on bi-encoders: models that compress each document into a single embedding vector ahead of time, then compare it to the query vector with a fast similarity search. That speed is why vector databases can search millions of records in under 100ms. But compressing an entire document into one vector loses information, so a genuinely relevant chunk can score just below your top_k cutoff and never make it into the LLM's context at all.
The fix is a two-stage pipeline. First, cast a wide net: retrieve a larger candidate set with vector search (say, the top 25 matches). Second, run a reranker, a cross-encoder model that reads the query and each candidate document together, in the same forward pass, instead of comparing precomputed vectors. Because it evaluates the actual query-document pair rather than two compressed summaries, it scores relevance far more precisely. A model like bge-reranker-v2-m3 is a common choice for this second pass.
In practice, this reordering step routinely promotes chunks that vector search buried at position 14 or 23 up into the top 3, the difference between the LLM seeing the right context and confidently answering from the wrong one. The tradeoff is cost: cross-encoders are too slow to run over your entire corpus, which is exactly why they're reserved for reranking a short candidate list rather than replacing vector search outright.
If your RAG answers feel inconsistent despite good source data, adding a reranking stage is often the single highest-leverage fix available.

References
Pinecone. (2026). Rerankers and two-stage retrieval. Pinecone Learning Center. https://www.pinecone.io/learn/series/rag/rerankers/