Advanced Retrieval: Techniques for When Simple Similarity Search Isn't Enough
Beyond basic vector search — HyDE, multi-query retrieval, fusion, contextual retrieval, graph-based retrieval, and multi-hop reasoning, and what specific retrieval failure each one is built to fix.
Introduction
Previous post Retreival Pipeline laid out the standard retrieval pipeline: process a query, embed it, retrieve candidates, rank, select, filter. That pipeline is enough for a large share of RAG use cases. But it has real, predictable failure modes — questions where the answer is scattered across multiple documents, questions phrased nothing like the way the answer is written, questions that require following a chain of relationships rather than matching a single passage. This post covers the techniques built specifically to address those gaps. Each one below exists because of a specific way plain retrieval breaks — worth keeping that framing in mind, since these aren’t upgrades you apply indiscriminately, they’re targeted fixes for targeted problems.
Query Expansion
Introduced briefly in retreival pipeline, query expansion earns a deeper look here because it’s the simplest and often first-reached-for advanced technique: rather than embedding a user’s query exactly as written, generate several related terms, synonyms, or rephrasings, and search using the expanded set rather than the original alone.
The problem this solves is a vocabulary mismatch between how a user asks a question and how the answer happens to be written in the source documents — a user asking about “cutting costs” might miss a document that only discusses “expense reduction,” even though a semantic embedding model narrows, but doesn’t fully close, this gap on its own. Expansion is typically done with an LLM prompted to generate a handful of alternative phrasings or related terms, which then get embedded alongside (or instead of) the original query, with results merged. It’s cheap, easy to add to an existing pipeline, and a reasonable first thing to try when retrieval seems to be missing relevant documents that share the same topic but different wording.
HyDE: Hypothetical Document Embeddings
HyDE takes a genuinely different, almost counterintuitive approach to the same underlying vocabulary-mismatch problem: instead of embedding the question, have an LLM first generate a hypothetical answer to the question, and embed that instead.
The reasoning behind this: a question and its real answer are often written in very different styles — a question is short, interrogative, and vague about specifics; a passage answering it is longer, declarative, and full of the exact terminology the question-writer didn’t know to use. By generating a plausible (even if not fully accurate) hypothetical answer first, HyDE produces text that’s stylistically and lexically much closer to how a genuine answer passage would actually be written — which means its embedding lands closer, in vector space, to real matching documents than the original terse question’s embedding would have.
The subtlety worth understanding: the hypothetical answer itself doesn’t need to be factually correct — it’s discarded after embedding, never shown to the user or used as the final answer. It only needs to be stylistically representative of what a real answer looks like, since that’s the only property being exploited. This makes HyDE surprisingly effective on exactly the kind of short, vague, or expertise-gapped queries where plain query embedding struggles most.
Multi-Query Retrieval
Multi-query retrieval generates several differently-phrased versions of the same underlying question — not synonyms for individual terms (that’s query expansion), but genuinely distinct full reformulations — runs retrieval independently for each one, and merges the combined candidate results.
The value here is a kind of ensemble effect: any single phrasing of a query has embedding “blind spots” — documents that happen to sit slightly outside where that particular phrasing’s vector lands, even though they’re genuinely relevant. Multiple differently-worded queries approach the same information need from different angles in embedding space, and merging their results substantially reduces the odds that a relevant document gets missed simply because of how one specific phrasing happened to embed. The cost is proportional — running k queries means k retrieval passes — which makes multi-query retrieval a good fit when answer completeness matters more than minimizing latency or compute cost.
Recursive Retrieval
Recursive retrieval structures retrieval as a multi-step process where each retrieval pass informs the next one, rather than firing a single query and treating whatever comes back as final. A common pattern: retrieve broadly first — perhaps at the document or section level — then run a second, more targeted retrieval pass restricted to just that narrowed subset, drilling down to the specific passage.
This mirrors the coarse-to-fine pattern that’s shown up repeatedly throughout this series — HNSW’s layered graph traversal, and the retrieve-then-rerank pattern— applied here at the level of retrieval strategy itself rather than inside a single algorithm or pipeline stage. It’s particularly useful for large, hierarchically organized knowledge bases, where jumping straight to fine-grained passage-level search across the entire collection is both slower and less accurate than first narrowing to the right general area.
Parent Document Retrieval
This is the retrieval-strategy counterpart to the parent-child chunking approach: search is performed against small, precise child chunks, but once a match is found, the larger parent chunk (or even the full source document) is what actually gets retrieved and passed to the LLM.
Worth being precise about the distinction from Chunking’s coverage: that post covered how you’d structure your chunks to support this pattern at ingestion time; this post’s framing is about the retrieval-time behavior that pattern enables — small-unit search precision, large-unit context delivery. It’s one of the more broadly applicable techniques in this post, since it doesn’t require any change to your query or an extra LLM call the way HyDE or query expansion do — the improvement comes purely from how your index and retrieval logic are structured.
Fusion Retrieval
Fusion retrieval is the general pattern of running multiple independent retrieval methods — commonly a semantic vector search alongside a lexical BM25 search, first introduced back in hybrid search discussion — and combining their separately-ranked result lists into one final ranking, rather than relying on any single method alone.
The standard combination technique is Reciprocal Rank Fusion (RRF), which scores each document based on its rank position in each individual result list (rather than the raw, differently-scaled similarity or BM25 scores themselves) and sums those rank-based scores across all the lists being fused. Using rank position rather than raw scores is the key design choice — it sidesteps the real practical problem that a BM25 relevance score and a cosine similarity score live on entirely different, incomparable scales, so naively averaging the raw numbers together would be mathematically meaningless. Fusion retrieval is essentially hybrid search generalized: nothing restricts it to exactly two methods — it’s a general mechanism for combining any number of independently-ranked retrieval passes into one list.
Contextual Retrieval
Contextual retrieval addresses a specific, easy-to-overlook chunking side effect: an isolated chunk, once split away from its source document, frequently loses context that a reader (or an embedding model) needs to understand what it’s actually about — a chunk might refer to “the company’s revenue that quarter” without ever stating which company or which quarter, because that information lived in an earlier paragraph the chunk boundary cut away from.
The technique, popularized by Anthropic’s own research on the topic, addresses this directly at ingestion time: before embedding each chunk, an LLM is given the chunk alongside its full source document and asked to generate a short, explanatory context — a sentence or two situating that specific chunk within the larger document — which is then prepended to the chunk before it’s embedded and indexed. This means the resulting embedding, and the text ultimately available to the retriever, carries context that would otherwise only have existed in the original, now-discarded document structure. It’s a genuinely different lever than the chunking strategies — rather than trying to draw smarter chunk boundaries, it accepts that some information loss at chunk boundaries is inevitable and compensates by re-injecting context directly into each chunk’s content.
Graph and Knowledge Graph Retrieval
Standard vector retrieval treats each chunk as an independent, isolated unit — but a great deal of real-world information is fundamentally relational: this person works at that company, this drug interacts with that medication, this clause references that other clause. Graph-based retrieval represents information as a network of entities and the relationships connecting them, rather than as a flat collection of independently-embedded text chunks.
A knowledge graph formalizes this as nodes (representing entities — people, organizations, concepts) connected by labeled edges (representing relationships between them — “works at,” “interacts with,” “is a subsidiary of”). Retrieval over a knowledge graph looks structurally different from vector similarity search: rather than finding the nearest chunk to a query embedding, it traverses relationships — starting from an entity mentioned in the query and following relevant edges outward to gather connected, relevant facts. This is a genuinely different retrieval paradigm from everything else in this series, better suited to questions that are fundamentally about how things relate to each other than to questions answerable from a single self-contained passage. Building and maintaining a knowledge graph is real, ongoing engineering effort — extracting entities and relationships accurately from unstructured text is its own non-trivial task — which is why graph retrieval tends to appear as a targeted addition for genuinely relational domains (organizational structures, regulatory frameworks, biomedical relationships) rather than as a default replacement for vector search.
Multi-Hop Retrieval
Multi-hop retrieval addresses questions that simply cannot be answered from any single retrieved passage, because the answer requires chaining together facts from multiple, separately-located sources — “What company did the founder of the company that acquired X previously work at?” requires first identifying X’s acquirer, then that acquirer’s founder, then that founder’s previous employer: three separate facts, three separate retrieval targets, each depending on the result of the one before it.
The mechanism: rather than a single retrieval pass, the system retrieves an initial fact, then uses that fact to formulate the next query, retrieves again, and repeats until enough connected information has been gathered to answer the original compound question — conceptually similar to recursive retrieval’s iterative structure, but driven specifically by genuine logical or factual dependency between hops rather than a coarse-to-fine narrowing of scope. This is one of the most computationally expensive techniques in this post — each hop is a full retrieval pass, often with an LLM call in between to determine the next query — which makes it worth reserving specifically for query patterns that are genuinely compositional, rather than applying it as a general-purpose upgrade to every query a system receives.
Closing Thoughts
None of the techniques in this post are meant to replace the standard retrieval pipeline — they’re targeted responses to specific ways that pipeline falls short: vocabulary mismatch (query expansion, HyDE), single-phrasing blind spots (multi-query), context lost at chunk boundaries (contextual retrieval, parent document retrieval), relational information (graph retrieval), and compositional questions (multi-hop). The skill worth building isn’t memorizing all of them — it’s recognizing which specific failure mode you’re actually looking at in your own retrieval results, so you reach for the technique built to fix that problem rather than layering on complexity that doesn’t address what’s actually going wrong.