What is Reranking?
Definition
Reranking is the step in which candidate documents or passages found by a fast first-stage retriever are scored again, and reordered, by a slower but more accurate model. The most common model for the job is a cross-encoder, which reads the query and the passage together. The goal is to make sure the few results shown to a user or passed to a language model really are the most relevant ones.
Also known as: re-ranking, reranker, cross-encoder reranking, second-stage ranking

The second stage of a two-stage search
Systems that search large collections strike a deal between speed and accuracy. First, a fast retriever narrows millions of documents down to, say, 100 candidates. Then a reranker reads only those 100, each alongside the query, scores them, and promotes the best five to ten. Running a slow, precise model over the whole collection would be impossible; running it over a shortlist is practical.
A reranker cannot recover a document the first stage missed. It only reorders what it is handed, which is why the two stages complement each other rather than compete.
Bi-encoders versus cross-encoders
| Bi-encoder (first stage) | Cross-encoder (reranking) | |
|---|---|---|
| Input | Query and document are encoded separately | Query and passage are read together as one input |
| Output | Two embeddings; similarity is computed afterwards | A single relevance score, directly |
| Precomputation | Document vectors are produced once, at indexing time | Not possible; the model runs again for every query-passage pair |
| Cost | Milliseconds across millions of documents | Grows linearly with the number of candidates |
| Accuracy | Good at topical similarity, can miss fine distinctions | Sees word-level relationships between query and passage directly |
The cross-encoder's edge comes from the transformer attention mechanism, which lets every word of the query interact with every word of the passage. For the query "how often should my cat be vaccinated?", a bi-encoder may score a passage about a dog's vaccination schedule highly, since both texts are about pet vaccines. A cross-encoder notices the mismatch between "cat" in the query and "dog" in the passage and pushes that passage down.
Other ways to rerank
- LLM rerankers. A language model receives the query and candidate passages and is asked either to score each one (pointwise) or to order the whole list (listwise). Flexible, but expensive and slow, and sensitive to how the prompt is written.
- Late interaction. Models such as ColBERT store token-level vectors for each document in advance and match them against the query at search time, a middle ground between bi-encoders and cross-encoders.
- Learning to rank. Classic search engines feed many features, such as text relevance, freshness and click data, into a trained ranking model.
Latency budgets and candidate counts
A cross-encoder's cost scales with the number of candidates: 100 candidates means 100 model passes. In practice two knobs are balanced. Sending more candidates to the reranker gives good results that the first stage ranked low a better chance of being rescued, but adds latency. Passing more reranked passages on to the language model gives the answer more evidence, but bloats the context and raises cost. The right values come from accuracy and latency targets measured on real queries. In systems that use hybrid search, the reranker usually runs after the two result lists have been fused.
What it means for content
A reranker reads each passage side by side with the query. A paragraph that is broadly on topic but never answers the question tends to drop at this stage; one that names the entities in the question and the relationship between them tends to rise. For a pricing query, "Our prices are competitive" will not match as well as "The enterprise plan is billed monthly per user, with a discount for annual payment." How passages become candidates in the first place is covered under passage retrieval, and the scoring logic behind the ranking under relevance scoring.

