In plain terms
The first search is a quick sweep of the shelves that pulls fifty possibly relevant books. Reranking is the careful step: someone reads the question next to each book and puts the truly useful ones on top. Only the top five go to the model, so the order matters.
Why it matters
Reranking is often the cheapest large improvement available to a RAG system: one extra step, no change to the index, and noticeably better passages in front of the model. It also lets you send fewer passages, which lowers cost and reduces the risk that irrelevant text distracts the model.
Example
A compliance assistant retrieves twenty passages for “Can we accept gifts from suppliers during a tender?”. The passage that answers it sits in ninth place, behind several that merely mention gifts. A reranker moves it to first. Sending the top four instead of the top twenty also cuts input tokens by eighty percent.
Most often confused with
Reranking vs. Retrieval
First-pass retrieval compares the query with millions of items using precomputed vectors, so it must be fast and is somewhat rough. Reranking looks at a few dozen candidates and reads query and passage together, which is far more accurate and far too slow to run on everything. Using them in sequence gets both speed and accuracy.
Under the hood
Rerankers are usually cross-encoders: the query and one passage are fed to a model together and a relevance score comes out, capturing interactions that separate embeddings miss. Alternatives include late-interaction models such as ColBERT and using an LLM to judge relevance. Typical setup: retrieve 50 to 200 candidates, rerank, keep 3 to 10. Costs: added latency of tens to hundreds of milliseconds and a per-query fee if hosted. A relevance threshold on the reranker score also lets the system say “nothing relevant found”, which is better than answering from weak passages.