Natural Language Processing

Reranking in RAG: Bi-Encoders, Cross-Encoders, and ColBERT Explained

Why a second, smarter ranking pass often matters more than the first retrieval pass — bi-encoders vs cross-encoders, ColBERT's late interaction, score fusion, and learned ranking models.

Introduction

Retrieval Pipeline introduced reranking briefly, as the refinement step sitting between candidate retrieval and final selection. This deserves its own in-depth understanding, because the architectural choice underneath reranking — how a model actually compares a query to a candidate — turns out to be one of the more consequential and least understood decisions in a RAG pipeline. This post is about that choice: the different ways a model can score query-document relevance, and why the fastest option and the most accurate option are, for a structural reason, never the same model.


Why Reranking?

Every retriever covered so far in this series — relies on comparing a query vector and a document vector that were each computed independently, then measuring the distance between them. That independence is exactly what makes retrieval fast enough to search millions of documents: document vectors get precomputed once, offline, and a query only needs to be compared against those already-finished vectors at search time.

But that same independence is also a real limitation. The two vectors were never actually computed with each other in mind — the model that embedded a document had no idea what query it would eventually be compared against, so it had to compress that document into a single fixed vector representing its meaning in general, not its relevance to any specific question. Reranking exists to add back exactly the thing this independence sacrifices: a scoring pass where the query and a specific candidate document are considered together, letting a model directly evaluate how well this particular document answers this particular question — at the cost of being far too slow to run across an entire collection, which is why it only ever runs on the smaller candidate set retrieval already narrowed down.


Bi-Encoders: The Architecture Behind Standard Retrieval

Every embedding-based retriever discussed in this series so far — without it being named explicitly — has been using a bi-encoder architecture: the query and the document are each passed through the model separately, producing two independent vectors, which are then compared using a similarity measure like cosine similarity.

The name reflects the structure: two (bi-) encoding passes that never interact with each other. This is precisely what makes bi-encoders fast at scale — since document vectors don’t depend on the query at all, they can be computed once, offline, and reused for every future query, and search becomes a matter of comparing a new query vector against a large set of already-finished vectors using the ANN algorithms. The trade-off is exactly the limitation named above: because the two vectors are computed blind to each other, a bi-encoder’s similarity score is a fixed, general-purpose approximation of relevance — it can’t adapt its judgment to the specific nuance of a particular query-document pairing the way a model that actually sees both together could.


Cross-Encoders: Precision at a Real Cost

A cross-encoder takes the opposite architectural approach: instead of encoding the query and document separately, it concatenates them into a single combined input and processes that pair together, in one pass, through the model — typically outputting a single relevance score directly, rather than two vectors to be compared afterward.

Because the model sees the query and document text simultaneously, it can attend to specific interactions between them — noticing, for instance, that a particular phrase in the document directly answers a particular clause in the question — in a way a bi-encoder’s independently-computed vectors structurally cannot. This consistently makes cross-encoders more accurate at judging relevance than bi-encoder similarity scores. The cost is equally structural and unavoidable: because the document half of the input depends on which query it’s being paired with, nothing can be precomputed offline — every single query-document pair requires its own full model pass at query time. This is why cross-encoders are never used for initial retrieval across an entire collection; they’re reserved specifically for reranking a small candidate set that a faster bi-encoder has already narrowed down.


ColBERT and Late Interaction

ColBERT (Contextualized Late Interaction over BERT) sits deliberately between the two extremes above, and understanding it clarifies what “late interaction” actually means as a design pattern in its own right.

Where a standard bi-encoder compresses an entire document into one single vector, ColBERT keeps a separate contextual embedding for every individual token in both the query and the document — a much richer, more granular representation. Relevance is then computed by comparing every query token embedding against every document token embedding, taking each query token’s best-matching document token, and summing those best-match scores together — a mechanism called MaxSim. Because these token-level embeddings are still computed independently for queries and documents (only the comparison step happens jointly, “late” in the process, rather than the encoding itself), document-token embeddings can still be precomputed and stored offline, just like a bi-encoder’s — the “late interaction” name refers precisely to this: interaction between query and document is delayed until the final comparison step, rather than happening during encoding as it does in a cross-encoder.

This gives ColBERT a genuinely useful middle position: meaningfully more accurate than a standard single-vector bi-encoder, because token-level matching captures finer-grained relevance signals a single compressed vector loses — while remaining far faster than a cross-encoder at query time, since the expensive part (encoding) is still precomputed offline. The trade-off moves to storage instead: keeping a vector per token, rather than one vector per document, multiplies the storage footprint substantially, which is the real cost ColBERT-style approaches are trading against their accuracy and speed gains.


Score Fusion

By this point in a RAG pipeline, you may have several independent relevance signals for the same candidate: a bi-encoder similarity score, a BM25 lexical score, and a cross-encoder or ColBERT reranking score. Score fusion is the general problem of combining multiple such signals into one final ranking, rather than picking just one and discarding the rest.

This connects directly back to the fusion retrieval technique — the same core challenge applies here: these scores live on entirely different, often incomparable scales (a cosine similarity between -1 and 1 has no natural relationship to a BM25 score or a cross-encoder’s raw logit output), so naively adding or averaging raw scores together tends to produce meaningless results dominated by whichever signal happens to have the largest numeric range. Rank-based fusion methods like Reciprocal Rank Fusion sidestep this by working with each signal’s rank ordering rather than its raw value, while more sophisticated approaches train a dedicated model specifically to learn how to weight and combine these signals — which leads directly into the final technique in this post.


Learned Ranking

Learning to Rank (LTR) treats the ranking problem itself as a supervised machine learning task: rather than hand-designing how to combine similarity scores, lexical scores, and other relevance signals, a model is trained directly on labeled data — queries paired with documents and human (or otherwise reliably established) relevance judgments — to learn the optimal way to combine whatever features are available into a final ranking.

These features can extend well beyond the retrieval-specific scores covered elsewhere in this post — document recency, source authority, historical click-through patterns, or any other signal correlated with genuine relevance in your specific domain can be incorporated directly into the model’s training data. This is a meaningfully heavier, more data-hungry approach than the other techniques — it requires a real labeled dataset of query-document relevance judgments to train against, which most teams building a RAG system don’t start out with — but for a mature system, especially one with real user interaction data (which queries led to which chunks actually being useful) to learn from, a trained ranking model can outperform any hand-tuned combination of the individual signals covered above.


Closing Thoughts

The throughline across every technique in this post is the same trade-off surfacing in a different shape each time: more interaction between query and document text produces more accurate relevance judgments, and always costs more compute or storage to achieve. Bi-encoders sacrifice interaction entirely for speed; cross-encoders embrace full interaction and sacrifice speed; ColBERT finds a genuine middle position by delaying interaction to the comparison step; score fusion and learned ranking are about combining whatever signals you already have as effectively as possible, rather than generating a new one. Knowing where a given technique sits on this trade-off is what lets you reason about reranking deliberately — matching the right amount of interaction, and the right cost, to what your latency budget and accuracy requirements actually demand.