Natural Language Processing

Chunking Strategy for RAG: The Decision That Quietly Determines Retrieval Quality

Why how you split documents matters as much as what embedding model you use — fixed, recursive, semantic, and structure-aware chunking strategies, plus how to actually choose chunk size and overlap.

Introduction

Last Post left us with clean, extracted text. Before that text can be embedded and stored in a vector database, it needs to be broken into smaller, retrievable pieces — chunks. This sounds like a mechanical detail. It isn’t. Chunking is one of the highest-leverage decisions in a RAG pipeline, and it’s also one of the most commonly underinvested in — teams will agonize over which embedding model to use, then split their documents with a single default parameter and move on. This post is about why that’s a mistake, and what the actual alternatives look like.


Why Chunking Matters

Two separate constraints force chunking to exist. First, a practical one: embedding models have a maximum input length, and even within that limit, embedding an entire long document into a single vector blurs together too many distinct ideas to be useful for precise retrieval. Second, a more subtle one: retrieval quality itself depends on chunks being the right size to answer a query without either drowning it in irrelevant surrounding text or truncating it before the useful part appears.

There’s a real tension underneath every chunking decision: chunks that are too large dilute relevance — a chunk covering five different subtopics might get retrieved for a query about just one of them, and then most of what the LLM receives as “context” is actually noise. Chunks that are too small lose context — a sentence pulled in isolation from its surrounding paragraph can lose the antecedent of a pronoun, the subject of a comparison, or the qualifying condition that made the original sentence true. Every chunking strategy below is, at its core, a different approach to managing this same trade-off.


Fixed-Size Chunking

The simplest strategy: split text into chunks of a fixed length — say, every 500 tokens — regardless of where sentences, paragraphs, or ideas happen to end.

Its appeal is obvious: it’s trivial to implement, fast, and completely predictable in terms of resulting chunk count and size. Its flaw is equally obvious once you see it: a fixed-length cut has no awareness of meaning, so it will just as happily slice a sentence in half, separate a heading from the content it introduces, or split a table’s header row from its data rows as it would find a clean paragraph boundary. Fixed-size chunking is a reasonable baseline to get a pipeline running end-to-end, but it’s rarely the right choice for a production system where retrieval quality actually matters.


Recursive Chunking

Recursive chunking improves on the fixed-size approach by respecting a hierarchy of natural boundaries rather than cutting blindly at a token count. The typical implementation tries to split on the largest natural boundary first — paragraph breaks — and only falls back to a smaller boundary (sentences, then words) if a resulting piece is still too large for the target chunk size.

This produces meaningfully cleaner chunks than fixed-size splitting, because it’s actively trying to avoid mid-sentence or mid-paragraph cuts wherever the text structure allows it, while still guaranteeing every chunk stays within your size limit. It’s a pragmatic middle ground — not aware of meaning the way semantic chunking (next) is, but aware of basic textual structure — and it’s become something close to a sensible default in many RAG frameworks for exactly that reason.


Semantic Chunking

Semantic chunking goes a step further: instead of relying on structural punctuation (paragraph breaks, sentence boundaries) as a proxy for where ideas begin and end, it uses embeddings themselves to detect actual topic shifts in the text.

A common implementation embeds individual sentences, then walks through the document measuring the similarity between consecutive sentences’ embeddings. A sharp drop in similarity between one sentence and the next is treated as a likely topic boundary — a natural place to start a new chunk, because the meaning of the text just changed direction, not merely because a certain number of tokens has elapsed. This tends to produce chunks that are more internally coherent — each one genuinely “about” one thing — at the cost of extra computation during ingestion (you’re now running embedding calls as part of chunking, not just at final indexing) and less predictable chunk sizes, since topic boundaries don’t respect a target token count.


Sliding Window Chunking

Sliding window chunking is less a distinct splitting strategy and more a modification applied on top of fixed-size or recursive chunking: instead of chunk boundaries sitting immediately adjacent to each other, consecutive chunks overlap by a set amount, so each new chunk “slides” forward by less than its full length rather than starting exactly where the previous one ended.

The motivation is direct: it reduces the chance that a piece of important context gets stranded right at a chunk boundary, split awkwardly between two chunks such that neither one contains the complete idea. We’ll return to exactly how much overlap is worth using later in this post, since it interacts directly with chunk size and comes with its own real costs, not just benefits.


Parent-Child Chunking

Parent-child chunking (sometimes called hierarchical chunking) resolves the earlier size trade-off with a genuinely different structural idea, rather than just tuning a single chunk-size number: instead of one chunk size for everything, it maintains two linked layers — small “child” chunks used for the actual similarity search, and larger “parent” chunks (often the full paragraph or section a child chunk belongs to) that get retrieved and handed to the LLM once a matching child chunk is found.

The reasoning: small chunks make for more precise retrieval — a tightly scoped chunk is easier to match confidently against a specific query — but small chunks also carry less context for the LLM to actually answer well. Parent-child chunking gets both properties at once: search happens against the precise, focused child chunks, but the LLM ultimately receives the richer parent chunk once a relevant child has been located. This pattern has become increasingly common in production RAG systems precisely because it sidesteps the small-vs-large trade-off rather than just picking a compromise point on it.


Structure-Aware Chunking

A cluster of related strategies fall under this umbrella, and they share one core idea: use whatever structural signals the source document already carries — the kind of structure Data Ingestion emphasized preserving during ingestion — as chunk boundaries, rather than treating the document as an undifferentiated stream of text.

Section-aware chunking uses a document’s actual heading hierarchy (the #, ##, ### structure from Markdown, or equivalent heading styles in DOCX and HTML) as natural chunk boundaries — keeping a section’s content together and rarely, if ever, needing to split mid-thought, since section boundaries were already chosen by whoever authored the document to represent genuine topic divisions.

Page-aware chunking respects page boundaries, most relevant for PDFs — useful less because a page break inherently signals a topic change (it usually doesn’t) and more because preserving page numbers in a chunk’s metadata lets you cite “page 12” back to a user, and because some documents (contracts, forms) are genuinely organized page-by-page in a way worth preserving.

Table-aware chunking treats tables as a distinct case entirely, rather than letting them get carved up by whatever generic chunking rule applies to surrounding prose. A table split across two chunks usually loses its header row in one half, making the data in that half meaningless in isolation — so table-aware strategies typically keep an entire table together as one chunk (or, for very large tables, repeat the header row in every sub-chunk) rather than applying the same token-count logic used for prose.

Figure-aware chunking handles images, charts, and diagrams — content that carries genuine information but isn’t text at all. The common approach is generating a text description of the figure (either through a caption already present in the source document, or a vision-capable model producing one during ingestion) and treating that description as its own retrievable chunk, linked back to the original image, so a query about a chart’s content has something textual to match against in the first place.

The unifying principle across all four: let the document’s own structure guide chunk boundaries wherever that structure is available, rather than defaulting to token-count-based splitting for content that already comes with a much better organizational signal built in.


Chunk Overlap

Returning to the sliding-window idea with more precision: overlap is typically expressed as a percentage of chunk size (say, 10–20%) or an absolute token count, and it exists to reduce the odds that a self-contained idea gets stranded across a chunk boundary with no single chunk containing the complete thought.

But overlap isn’t free, and it’s worth being deliberate about the trade-off rather than defaulting to a large value out of caution. More overlap means more total chunks for the same source text, which means more vectors to store, more compute spent embedding redundant content, and — somewhat counterintuitively — a real risk of the same underlying content being retrieved multiple times as separate near-duplicate chunks, crowding out other genuinely distinct information from the LLM’s limited context window. A modest overlap (commonly in the 10–15% range) is usually enough to catch boundary-split ideas without meaningfully bloating the index; very large overlap values tend to trade a marginal robustness gain for a real cost in redundancy.


Chunk Size Optimization

There’s no universally correct chunk size — it’s a genuine trade-off that depends on your content and your queries, and it’s worth arriving at through testing rather than adopting a default number uncritically.

Smaller chunks (roughly 100–300 tokens) tend to produce more precise retrieval — the embedding represents a narrow, specific idea, so similarity matching against a specific query is sharper — but they carry less context individually, which matters more if you’re not using a parent-child strategy to compensate. Larger chunks (500–1000+ tokens) preserve more surrounding context per chunk, reducing the odds a critical qualifying detail got separated from the fact it modifies, but dilute the embedding’s specificity and make it easier for irrelevant content to get pulled in alongside anything genuinely relevant.

The right way to actually settle this, rather than guessing: build a small evaluation set (the same recommendation from post of Embeddings and Embedding Models, applied here) with representative queries and their known-correct source passages, then test a handful of chunk-size and overlap configurations against it, measuring retrieval metrics like recall@k directly — not just eyeballing whether a few example queries happen to return something plausible. Chunk size interacts with your document type too: dense technical documentation and long narrative prose genuinely warrant different chunk sizes, so a single project may reasonably need more than one chunking configuration if it ingests meaningfully different kinds of source content.


Closing Thoughts

Chunking sits at the exact intersection of everything covered so far in this series — it determines what actually gets embedded, what gets stored and searched in the vector database, and it’s shaped directly by how well ingestion preserved document structure. Getting it right is rarely about picking the single “best” strategy in the abstract — it’s about matching the strategy to your content’s actual structure and testing chunk size empirically, the same disciplined approach this series has returned to again and again.

With documents now cleanly ingested and properly chunked, next we will turn to what happens on the other side of retrieval: how a retrieved chunk actually gets combined with a user’s query and handed to an LLM to generate a grounded, accurate answer.