Natural Language Processing

The Retrieval Pipeline: How a Question Becomes the Right Context

Tracing the full path from a raw user query to the exact chunks handed to an LLM — query processing, candidate retrieval, ranking, top-k selection, and metadata filtering working together.

Introduction

Every post so far has built one piece of the machine: embeddings turn text into vectors, specific models produce those vectors, vector databases store them at scale using ANN algorithms to search them fast, and clean, well-chunked documents are what actually gets stored. This post assembles those pieces into the thing that runs every time a user asks a question: the retrieval pipeline — the sequence of steps between “here’s a question” and “here’s the exact context an LLM should answer it with.”


The Retriever as a Pipeline, Not a Single Step

It’s tempting to think of “retrieval” as one atomic operation — send a query, get back documents. In any production RAG system, it’s actually a sequence of distinct stages, each of which can be built well or built poorly, and each of which independently affects final answer quality. Treating retrieval as a single black box is exactly why so many RAG systems are hard to debug: when the wrong context reaches the LLM, there are several genuinely different places the failure could have occurred, and each needs a different fix.

The stages, in order, are: the raw query gets processed and possibly reformulated, it’s converted into an embedding, that embedding is used to pull a set of candidate matches from the vector database, those candidates are ranked (and sometimes reranked with a more expensive, more accurate model), a final subset is selected to actually send forward, and metadata filters are applied somewhere in that sequence to enforce hard constraints similarity search alone can’t express. The rest of this post walks through each stage in that order.


Query Processing

The user’s raw question is rarely the ideal thing to search with directly, and query processing is the stage that closes that gap before anything gets embedded.

Real user queries carry problems a retriever has to work around: they can be ambiguous (“what about the new one?” only makes sense with conversation history), underspecified (a single keyword instead of a full question), or phrased in a way that doesn’t semantically resemble how the answer is actually written in the source documents. Common query processing techniques address these directly: query rewriting uses an LLM to reformulate a vague or context-dependent question into a self-contained, explicit one before it’s embedded; query expansion adds related terms or rephrasings to broaden what the search can match against; and query decomposition breaks a genuinely multi-part question (“compare X and Y’s approach to Z”) into separate sub-queries that get retrieved independently and combined afterward, since a single embedding often can’t represent a compound question well.

This stage is easy to skip when a RAG system is first prototyped — the raw query gets embedded directly, and it often works fine on simple, well-phrased test questions. It tends to become necessary precisely when real users start asking the kind of messy, contextual, multi-part questions people actually ask.


Embedding the Query

Once the query is in its final, processed form, it’s converted into a vector using the same embedding model used to embed the document collection — a requirement, not a preference, since (as discussion of embedding drift established) comparing vectors from two different embedding spaces produces meaningless distances.

Recall from Embedding Models that some models are trained asymmetrically, expecting an explicit "query: " versus "passage: " prefix, or use structurally different encoding paths for queries versus documents. Getting this detail wrong — embedding a query the same way you’d embed a passage, when the model expects otherwise — is a subtle, easy-to-miss mistake that silently degrades retrieval quality without throwing any error, since the embedding call still succeeds; it just produces a vector that isn’t optimally positioned relative to your document embeddings.


Candidate Retrieval

This is the stage most people picture when they hear “retrieval”: the query embedding is handed to the vector database’s ANN index and a set of candidate matches — typically far more than what will ultimately be used — comes back, ordered by initial similarity score.

The key word here is candidate. This first pass is deliberately over-inclusive — retrieving, say, the top 50 or 100 nearest matches rather than only the handful that will actually be sent to the LLM — because the fast ANN similarity score used at this stage is a reasonable but imperfect proxy for true relevance. Casting a slightly wider net here, then narrowing it down more carefully in the ranking stage that follows, tends to produce better final results than trying to get everything right in a single retrieval pass.

Many production systems also perform hybrid retrieval at this stage — running both the semantic vector search and a lexical search like BM25 in parallel, then merging their candidate lists — for exactly the reasons covered back when we first introduced hybrid search: the two approaches fail in different, largely non-overlapping ways.


Ranking

Candidate retrieval produces a rough, fast ordering; ranking refines it. The simplest version of this stage is just sorting candidates by their similarity score from the previous step — but production systems frequently add a reranking step: passing the query alongside each candidate chunk through a separate, more powerful (and more computationally expensive) model specifically trained to judge relevance directly, rather than relying purely on embedding-space distance.

The reason this second pass helps: the fast ANN retrieval stage is optimized for speed across the whole collection, using a relatively coarse similarity signal. A reranker, by contrast, only has to evaluate the much smaller candidate set that survived the first stage, so it can afford to use a slower, more accurate model — often a cross-encoder, which processes the query and a candidate chunk together in a single pass (rather than comparing two independently-computed vectors), letting it directly model interactions between the two texts that a similarity score computed after the fact simply can’t capture. This two-stage pattern — fast, wide retrieval followed by slower, narrow reranking — is a close cousin of the same coarse-to-fine idea underlying HNSW’s layered graph traversal, just applied at the pipeline level instead of inside a single index.


Selecting What to Reach: Top-K and Threshold Retrieval

After ranking, a decision has to be made about exactly how much context actually gets forwarded to the LLM — and there are two structurally different ways to make that cut.

Top-K selection takes a fixed number of the highest-ranked results, regardless of how good or poor their absolute relevance actually is — if you ask for the top 5, you get 5, even on a query where only two results are genuinely relevant and the rest are mediocre at best. It’s simple, predictable, and keeps the LLM’s context size consistent across queries, which makes cost and latency easier to reason about.

Threshold retrieval instead sets a minimum relevance score and returns however many results clear that bar — which could be zero, or could be far more than a fixed top-k would have allowed. This avoids handing the LLM low-quality, barely-relevant context just to hit a fixed count, but it introduces its own real failure mode: if your knowledge base genuinely doesn’t contain a good answer to a query, threshold retrieval can correctly return nothing — which is the right behavior, but only if the rest of your system is built to handle an empty retrieval result gracefully (explicitly telling the user no relevant information was found) rather than silently prompting the LLM with no context and letting it guess anyway.

Many production systems combine both: a threshold to filter out clearly irrelevant candidates, and a top-k cap on top of that as a safety limit on context size — getting threshold retrieval’s quality floor together with top-k’s predictability ceiling.


Metadata Filtering

The final piece, threaded through the pipeline rather than confined to one stage: metadata filtering (introduced in Vector DB’s ) applies hard, structured constraints that similarity search alone has no mechanism to express — restricting candidates to a specific date range, document type, access-permission level, or department, for instance.

Where in the pipeline this filter gets applied matters, echoing the pre-filtering versus post-filtering distinction from vector DB’s post: applying it before candidate retrieval (so the ANN search only considers already-eligible vectors) is generally more reliable than applying it after ranking, which risks discarding good candidates simply because they were filtered out downstream of an otherwise-correct search, having taken up space in the initial results that should have gone to a genuinely matching candidate. For access-control-sensitive metadata specifically — who’s allowed to see a given chunk — pre-filtering isn’t just a performance optimization, it’s the only version of filtering that reliably prevents an unauthorized chunk from ever entering the ranking or generation stage at all.


Closing Thoughts

A retrieval pipeline is best understood as a funnel: broad, fast candidate retrieval narrows through progressively more careful and more expensive stages — ranking, reranking, selection, filtering — until only the most relevant, permitted, appropriately-sized context remains. When a RAG system gives a bad answer, this post’s stage-by-stage breakdown is the actual debugging checklist: was the query processed well, embedded correctly, given enough candidates to work with, ranked accurately, cut off at the right point, and filtered correctly — rather than treating “retrieval” as one indivisible thing to blame.