Knowledge & retrieval

Reranking

A second pass that re-orders retrieved passages by how well each one answers the query, so that the best few are the ones the model receives.

First-pass retrievalfast but rough ranking1 · Shipping and delivery times2 · Membership benefits3 · Steps to cancel a subscriptionRerankingreads each candidate with the query1 · Steps to cancel a subscription2 · Membership benefits3 · Shipping and delivery timesQuery: “how do I cancel my subscription?” The right passage moves from third place to first.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/reranking

In plain terms

The first search is a quick sweep of the shelves that pulls fifty possibly relevant books. Reranking is the careful step: someone reads the question next to each book and puts the truly useful ones on top. Only the top five go to the model, so the order matters.

Why it matters

Reranking is often the cheapest large improvement available to a RAG system: one extra step, no change to the index, and noticeably better passages in front of the model. It also lets you send fewer passages, which lowers cost and reduces the risk that irrelevant text distracts the model.

Example

A compliance assistant retrieves twenty passages for “Can we accept gifts from suppliers during a tender?”. The passage that answers it sits in ninth place, behind several that merely mention gifts. A reranker moves it to first. Sending the top four instead of the top twenty also cuts input tokens by eighty percent.

Most often confused with

Reranking vs. Retrieval

RerankingOrders a small set carefully
RetrievalFinds a broad set quickly

First-pass retrieval compares the query with millions of items using precomputed vectors, so it must be fast and is somewhat rough. Reranking looks at a few dozen candidates and reads query and passage together, which is far more accurate and far too slow to run on everything. Using them in sequence gets both speed and accuracy.

Under the hood

Rerankers are usually cross-encoders: the query and one passage are fed to a model together and a relevance score comes out, capturing interactions that separate embeddings miss. Alternatives include late-interaction models such as ColBERT and using an LLM to judge relevance. Typical setup: retrieve 50 to 200 candidates, rerank, keep 3 to 10. Costs: added latency of tens to hundreds of milliseconds and a per-query fee if hosted. A relevance threshold on the reranker score also lets the system say “nothing relevant found”, which is better than answering from weak passages.

Written by Mehmet Erkek · Last updated: