RAG & Knowledge Systems · 6 min
Reranking with Cross-Encoders: Precision on the Shortlist
First-stage retrieval encodes query and document apart for speed; a cross-encoder reads them together to fix the ranking on the top-k.
Your vector search returns 100 passages. Most are roughly on-topic, but the truly correct one sits at rank 14, and the user only reads the top 3. A cross-encoder reranker exists to close exactly that gap: pull the right passage from 14 to 1. Doing it well means understanding why the retriever put it at 14 in the first place.
Two ways to score a (query, document) pair
First-stage retrieval uses a bi-encoder. Query and document go through the encoder separately, and each collapses to a single fixed vector. Relevance is the cosine similarity between those two vectors. The thing that makes this fast is that the document never sees the query. You can encode all 8.8 million passages offline, store them in an index, and at query time only encode the query and do an approximate-nearest-neighbor lookup. Millions of documents, single-digit milliseconds.
The price is that everything the model knows about a document has to survive being crushed into one vector before it knows what you asked. The document vector for a passage about "jaguar mating range" is fixed whether your query is about the cat or the car. Similarity in that space is a decent proxy for relevance, not relevance itself.
A cross-encoder throws out the separation. It concatenates query and document into one sequence and runs the full transformer over the pair, so every query token attends to every document token through all layers. There is no reusable document vector. The score is a function of the specific pair.
BI-ENCODER (retrieval) CROSS-ENCODER (reranking)
query -> [enc] -> q vec [CLS] query [SEP] document [SEP]
doc -> [enc] -> d vec (offline) |
\ / [transformer: q<->d attention]
cos(q, d) |
[CLS] -> linear -> relevance score
precompute docs once, nothing precomputable;
score = cheap dot product one full forward pass PER pairThat joint attention is where the accuracy comes from. The model can represent "this passage answers this question" because it reads them together: negation, the entity you actually meant, whether the numbers in the passage match the numbers in the query. It is also why you cannot run it over the corpus. There is nothing to precompute, so scoring 8.8M passages per query means 8.8M forward passes. That is not slow, it is infeasible.
So you compose the two. Bi-encoder for recall (get the right passage somewhere in the top-k cheaply), cross-encoder for precision (order that shortlist correctly). This is the same recall-then-precision funnel you met in hybrid retrieval, where BM25 and dense vectors are unioned to widen recall, with the reranker bolted on as a precision stage after the union.
Where this came from: monoBERT
The template is Nogueira and Cho's 2019 paper "Passage Re-ranking with BERT," later dubbed monoBERT. The construction is almost aggressively simple. Feed BERT [CLS] query [SEP] passage [SEP], take the [CLS] output vector, push it through one linear layer to a single relevance score, and train with binary cross-entropy on relevant versus non-relevant pairs. At inference you score each candidate independently and sort by the score.
Dropped onto the MS MARCO passage leaderboard, it beat the prior state of the art by 27% relative in MRR@10, a large jump for a method that fits in a paragraph. The follow-up line matters too. duoBERT made scoring pairwise ("is A more relevant than B?"). monoT5 (Nogueira et al., 2020) reframed the whole thing as generation: feed a T5 the query and document, ask it to emit the token true or false, and rank by the logit of true. Same joint-attention idea, different head. That is the ancestry behind every cross-encoder/* reranker you pip-install today, and behind the current LLM listwise rerankers (RankT5, RankZephyr) that read the whole shortlist at once instead of scoring pairs independently.
The numbers that decide your design
The public MS MARCO cross-encoders make the trade-off concrete (Sentence-Transformers benchmark, single GPU):
| Model | NDCG@10 (TREC DL19) | Docs/sec |
|---|---|---|
| ms-marco-TinyBERT-L2-v2 | 69.8 | ~9000 |
| ms-marco-MiniLM-L6-v2 | 74.3 | ~1800 |
| ms-marco-MiniLM-L12-v2 | 74.3 | ~960 |
Read that table as a budget, not a leaderboard. "Docs/sec" is candidates scored per second on one GPU. Rerank the top 100 with the L6 model and you spend about 55ms of GPU time per query before batching. Rerank the top 1000 and you are at roughly 550ms, often slower than the retrieval it is correcting. Notice also that L6 and L12 tie on quality here. Past a point, a bigger reranker buys latency, not accuracy. The usual sweet spot is retrieve 50 to 200, rerank with an L6-class model, keep the top 3 to 10.
Builder: the reranker only re-orders what retrieval already found. If the right passage is not in your top-k, no reranker recovers it. That is a recall miss, and the fix is upstream (chunking, hybrid retrieval, a better embedding model), not a stronger cross-encoder. Measure recall@k of your first stage before you tune the reranker, or you will be sanding a board that isn't there.
What actually goes wrong
Silent truncation. Cross-encoders have a fixed max sequence length, often 512 tokens for query + document combined. Feed a chunk longer than that and the tail is cut before scoring. The model rates a passage on evidence it never saw, and a relevant sentence at the bottom is invisible to it. Log token lengths and size chunks to the reranker's window.
Uncalibrated scores. monoBERT-style logits are trained for ranking, not probability. A score of 8.2 for one query and 3.1 for another tell you nothing comparable, so a global "relevance > 0.5" cutoff will not mean the same thing across queries. Threshold within a query (keep the top-k, or watch for a relative drop-off), not against an absolute constant.
Latency under load. That 55ms is single-query. Fan-out (reranking for every sub-query of a decomposed question) or a traffic spike multiplies GPU passes with no precompute to amortize. Cache (query, doc-id) to score, batch aggressively, and cap k.
Defender: the reranker reads attacker-controllable text, the retrieved documents, through a full transformer that attends to the query. A poisoned chunk stuffed with your exact query phrasing, or with imperative text aimed at a downstream generator, can ride a high rerank score straight to the top of context. Reranking is a relevance filter, not a safety filter, and it will happily promote a hostile passage that looks maximally on-topic. Sanitize and provenance-check after reranking; never assume rank 1 is trustworthy.
Researcher: the monoBERT to monoT5 to listwise-LLM progression is a clean testbed. Pairwise (duoBERT) and listwise rerankers beat pointwise because scoring candidates against each other captures relative relevance that independent scoring misses, at O(k²) or full-context cost. If you evaluate rerankers, report first-stage recall@k alongside end NDCG. A reranker's ceiling is set entirely by what retrieval handed it, and papers that omit that number are measuring two things at once.
The mental model to keep
The bi-encoder answers "which documents are plausibly about this?" cheaply and at corpus scale. The cross-encoder answers "does this document answer this query?" expensively, on a shortlist. You need both because neither scales into the other's job: joint attention does not fit over millions of docs, and separated vectors cannot capture true pairwise relevance. It is the same funnel as hybrid retrieval, widen recall then narrow precision, with the reranker as the last and sharpest narrowing before text reaches the generator.
Sources
- Nogueira & Cho, "Passage Re-ranking with BERT," 2019 — arXiv:1901.04085 (monoBERT; code at github.com/nyu-dl/dl4marco-bert)
- Nogueira, Jiang, Pradeep & Lin, "Document Ranking with a Pretrained Sequence-to-Sequence Model," Findings of EMNLP 2020 — aclanthology.org/2020.findings-emnlp.63 (monoT5) — https://aclanthology.org/2020.findings-emnlp.63/
- Sentence-Transformers docs — Cross-Encoder pretrained models (MS MARCO NDCG@10 / docs-per-sec table) and the "Retrieve & Re-Rank" guide, sbert.net — https://www.sbert.net/docs/pretrained-models/ce-msmarco.html
- Lin, Nogueira & Yates, "Pretrained Transformers for Text Ranking: BERT and Beyond," 2020 — arXiv:2010.06467 (mono/duo taxonomy, multi-stage ranking)
- MS MARCO passage ranking dataset (8.8M passages) — microsoft.github.io/msmarco — https://microsoft.github.io/msmarco/