The End-to-End RAG Pipeline: How Every Piece Fits Together
Assembling fourteen chapters into one coherent system — the offline indexing path, the online query path, and the error handling that separates a working demo from something production-ready.
Introduction
Posts so far, have covered a piece in isolation: embeddings, vector databases, ANN algorithms, chunking, retrieval, reranking, prompting, LLM integration. This post assembles them. Not as a summary — as an architecture. Because the thing that trips up most teams building their first production RAG system isn’t understanding any individual component; it’s understanding how the components relate, where the boundaries between them sit, and which parts run when.
The Two-Path Architecture
The single most useful structural insight about RAG systems is this: a RAG pipeline is not one pipeline. It’s two, running at completely different times, at completely different frequencies, with completely different performance requirements.
The indexing path (sometimes called the ingestion or offline path) runs ahead of time, independent of any user. It takes raw source documents and transforms them into a searchable index. It runs periodically — on a schedule, or triggered when source data changes — and it can afford to be slow, because no user is waiting on it.
The query path (the online or runtime path) runs every time a user asks a question. It takes a query, finds relevant context in the index the other path built, and generates an answer. It has to be fast, because a user is very much waiting on it.
INDEXING PATH (offline, periodic)
Sources → Parse → Clean → Chunk → Embed → Store in Vector DB
│
▼
[ Vector Index ]
│
▼
QUERY PATH (online, per request)
Query → Process → Embed → Retrieve → Rerank → Prompt → Generate → Respond
Almost every architectural decision in a RAG system becomes clearer once you know which path it belongs to. Expensive, high-quality operations (semantic chunking, contextual retrieval’s per-chunk LLM summaries, careful cleaning) belong on the indexing path where latency doesn’t matter. Fast approximations (ANN search, top-k truncation) belong on the query path where it does.
Data Flow
Tracing a single piece of information through both paths makes the boundaries concrete.
A paragraph in a PDF starts as pixels-and-layout instructions inside a file. The indexing path extracts it as text, cleans away headers and broken whitespace, splits it into a chunk with its neighbors, converts that chunk into a vector via an embedding model, and stores that vector alongside its original text and metadata in a vector database. At this point, the paragraph exists in the system as a searchable unit — and nothing has happened yet from any user’s perspective.
Later, a user asks a question. The query path converts that question into a vector using the same embedding model, searches the index via an ANN algorithm to find the nearest chunk vectors, reranks the candidates with a more careful model, assembles the surviving chunks into a prompt within a token budget, and sends it to an LLM. If that original paragraph was relevant, its text — not its vector — is what reaches the model.
The detail worth internalizing: vectors are how you find content; text is what you actually use. The embedding exists purely to make the paragraph findable. Once found, the original text is what gets injected into the prompt. A surprising number of conceptual confusions about RAG dissolve once this distinction is clear.
The Indexing Pipeline
The offline path, stage by stage, with the post each stage draws on:
- Source connection — pulling documents from files, databases, APIs, or SaaS platforms, including the sync strategy that keeps them current.
- Parsing and extraction — converting each format into clean text, handling PDFs, OCR, HTML boilerplate, and tabular data.
- Cleaning and normalization — removing boilerplate, fixing encoding, de-duplicating.
- Chunking — splitting text into retrievable units using a strategy matched to the content’s structure.
- Enrichment — optionally attaching metadata and generated context to each chunk.
- Embedding — converting chunks to vectors in batches.
- Storage and indexing — writing vectors, text, and metadata to the vector database and building its ANN index.
Two operational properties matter enormously here and are easy to skip in a prototype. First, incrementality: re-indexing an entire corpus every time one document changes is wasteful and slow at real scale, so production pipelines track which sources changed and update only those — which requires stable chunk IDs that survive re-processing. Second, idempotency: running the pipeline twice on the same unchanged document should produce the same result, not duplicate entries, which is what makes retries safe when a run fails partway through.
The Retrieval Pipeline
Here it’s worth seeing where it sits relative to everything around it. Retrieval begins after a query arrives and ends when a final set of chunks has been selected — it produces context, not an answer.
The stages: query processing and optional reformulation, embedding the processed query, candidate retrieval from the ANN index (usually over-fetching deliberately), reranking the candidates with a cross-encoder or similar, applying metadata filters, and selecting a final set via top-k, threshold, or both.
The boundary matters architecturally: because retrieval outputs context rather than prose, it can be tested, measured, and improved entirely independently of the LLM. You can compute recall@k on a retrieval pipeline without generating a single answer — which is exactly what makes debugging tractable, since it lets you isolate whether a bad answer came from bad context or bad generation.
The Generation Pipeline
Generation takes selected context plus the user’s question and produces an answer. Its stages: assembling the prompt from a template, fitting everything within the token budget, calling the model (streaming or not), and receiving the output.
The critical design decision here is what happens when retrieval returns little or nothing useful. A pipeline that unconditionally hands whatever it got to the LLM and asks for an answer will get one — fluent, confident, and unsupported. A well-built generation stage checks first: if no chunk cleared the relevance threshold, the correct behavior is usually to short-circuit and tell the user the knowledge base doesn’t cover this, rather than prompting a model with empty or marginal context and hoping the system prompt’s grounding instructions hold. This is where threshold retrieval and abstention instructions meet as one concrete architectural branch.
The Response Pipeline
What happens between the model finishing and the user seeing something is its own stage, and it’s the one most often left out of architecture diagrams entirely.
Typical responsibilities: post-processing the raw output (parsing structured output, resolving citation markers back to real source documents with titles and links), validation (checking that cited chunks actually support the claims attached to them, running faithfulness checks), presentation (formatting the answer, attaching source references the user can click through to), and logging (recording the query, retrieved chunk IDs, scores, and the final answer).
That last one deserves emphasis. Logging the full retrieval trace — not just the final answer, but which chunks were retrieved, with what scores, and which survived reranking — is what makes a production RAG system debuggable at all. Without it, a user reporting a bad answer gives you nothing to investigate. With it, you can replay exactly what the system saw and pinpoint which stage failed. It also accumulates into the labeled data that makes learned ranking and evaluation possible later.
Error Handling
Every stage above can fail, and RAG systems have an unusual property that makes error handling genuinely tricky: most failures are silent. A crashed API call is obvious. A retrieval pass that returned plausible-but-wrong chunks produces a fluent, confident, entirely incorrect answer with no error anywhere in the logs.
Failures worth designing for explicitly:
Hard failures— the vector database is unreachable, the embedding API times out, the LLM call is rate-limited. These need retries with backoff, timeouts, and a graceful user-facing message rather than a stack trace.Empty retrieval— no chunk cleared the threshold. Not an error, but a distinct path requiring its own handling, as covered above.Degraded retrieval— chunks were returned, but with uniformly low scores. Worth surfacing to the user as low confidence rather than silently treating it as a normal result.Context overflow— retrieved chunks exceed the token budget. Should truncate deliberately by rank (dropping lowest-ranked chunks) rather than failing or letting the model silently lose the end of its context.Ingestion failures— a PDF that fails to parse, a document that produces zero chunks. These are easy to miss because the pipeline “succeeds” overall; a document that silently never made it into the index is indistinguishable, at query time, from one that simply doesn’t contain the answer.
The design principle underneath all of these: a RAG system should fail loudly upstream rather than quietly downstream. A parse failure caught during indexing is a fixable ticket. The same failure discovered because a user got a confidently wrong answer three months later is a trust problem — and by then, you have no obvious way to know it was a parsing issue at all.
Closing Thoughts
The architecture in this post is deliberately conventional, because a conventional RAG pipeline built carefully outperforms an exotic one built carelessly. The two-path split, clean boundaries between stages, honest handling of empty and degraded results, and logging thorough enough to reconstruct what happened — these are what make a system maintainable and improvable rather than a black box that works until it doesn’t.
Which raises the question this series has deferred several times now: how do you actually know it’s working? Next we take up on evaluation — the metrics, datasets, and methods for measuring RAG quality end to end, and for catching regressions before your users find them.