Natural Language Processing

Data Ingestion for RAG: Turning Messy Real-World Documents Into Usable Text

The unglamorous but decisive first step of any RAG pipeline — parsing PDFs, OCR, HTML, DOCX, databases, and APIs, and why data cleaning quietly determines how good your retrieval can ever be.

Introduction

Every post in this series so far has assumed clean text already exists somewhere, ready to be embedded and indexed. In reality, that text starts out trapped inside PDFs with broken layouts, scanned images, HTML pages full of navigation clutter, and database rows that were never meant to be read as prose. Data ingestion is the unglamorous work of turning all of that into usable text — and it deserves far more respect than it usually gets, because no embedding model or retrieval algorithm from earlier posts can rescue a RAG system fed badly extracted text. This is the “garbage in, garbage out” stage, and it’s where a surprising share of real-world RAG problems actually originate.


The Many Shapes of Source Data

Before touching any parsing tool, it’s worth categorizing what you’re actually dealing with, because the right ingestion strategy depends entirely on the source format.

Source data broadly falls into a few buckets: unstructured text (plain prose with no explicit schema — web pages, Word documents, emails), semi-structured text (has some organizational structure but isn’t rigidly tabular — Markdown, HTML, JSON), structured data (rigidly organized into rows and columns — CSVs, database tables), and non-text sources (images, scanned documents, audio transcripts) that need an extra conversion step before they’re text at all.

The mistake worth avoiding early: treating every source format with the same generic “extract all text” approach. A PDF, an HTML page, and a database table each carry different structural signals — headings, tables, rows — that are worth preserving deliberately, not stripping away, because that structure often carries meaning your retrieval system will later depend on.


Parsing PDFs

PDFs are, without much competition, the most painful common ingestion source in RAG pipelines — worth understanding why, since it shapes what “parsing” actually has to accomplish.

A PDF is fundamentally a layout format, not a text format — it specifies where characters should be drawn on a page, not the logical reading order or structure of the content. This creates real problems: a two-column academic paper can have text extracted in a jumbled, interleaved order if a parser naively reads left-to-right across the whole page width. Tables are especially fragile — what looks like a clean grid visually is often just a collection of individually positioned text fragments with no inherent “this is a table” markup, so naive extraction can turn a table into a scrambled wall of numbers with no columns.

Modern PDF parsing libraries address this with layout-aware extraction — detecting columns, tables, and reading order using the visual positioning of text rather than just its raw stream order — and some incorporate machine learning models specifically trained to recognize document structure (headers, tables, figures, footnotes) rather than treating a page as an undifferentiated blob of characters. The practical takeaway: always spot-check extracted PDF text against the original visually, especially for anything with tables, multi-column layouts, or footnotes — silent extraction errors here are common and easy to miss until retrieval quality mysteriously suffers later.


OCR: Reading Text from Images

Some PDFs — and standalone image files — contain no extractable text at all, because they’re scans: a photograph or scanned image of a page, with no underlying text layer, just pixels. Optical Character Recognition (OCR) is the technology that converts these pixels back into machine-readable text.

Modern OCR systems, particularly ones built on deep learning rather than older template-matching approaches, handle a wide range of fonts, languages, and moderate image quality issues well — but OCR quality still degrades meaningfully with poor scan resolution, skewed or rotated pages, handwriting, or unusual layouts (forms, tables, multi-language documents). Every OCR error becomes a factual error baked directly into your knowledge base — a misread digit in a scanned financial table, for instance, becomes a confidently wrong fact your RAG system will retrieve and present as ground truth. For any RAG system whose accuracy actually matters, validating OCR output — even spot-checking a sample against source images — is worth the effort it takes.


Parsing HTML

Web pages carry a structural advantage PDFs don’t: explicit markup (<h1>, <table>, <p>, and so on) that directly signals document structure — which makes HTML meaningfully easier to parse correctly, if you use that structure rather than discarding it.

The real challenge with HTML isn’t extracting the text — it’s separating the content that actually matters from the substantial amount of surrounding clutter modern web pages carry: navigation menus, sidebars, footers, cookie banners, ads, and related-article widgets. Naive extraction (stripping all HTML tags and keeping whatever text remains) tends to pull in this boilerplate alongside the genuine content, diluting a retrieval chunk with irrelevant text. Better approaches use readability-style extraction — heuristics or models specifically designed to identify a page’s main content region, the same underlying technique your browser’s “reader mode” uses — to isolate the article or documentation content from everything surrounding it.


Parsing Markdown

Markdown is, in a real sense, the easiest ingestion source covered in this post — it’s plain text with lightweight, unambiguous structural markers (# for headings, - for lists, and so on) designed from the outset to be both human-readable and simple to parse programmatically.

The main thing worth doing deliberately with Markdown, rather than treating it as an afterthought, is preserving its heading hierarchy during ingestion rather than flattening it into undifferentiated prose. A document’s #, ##, and ### structure is a strong, explicit signal about topic boundaries and hierarchy — information that’s directly useful for smarter chunking strategies later, so it’s worth carrying that structure forward into your metadata rather than discarding it the moment the text is extracted.


Parsing DOCX

Word documents are structured XML under the hood (a .docx file is technically a zip archive of XML files), which makes them more machine-parseable than PDFs in principle — but they carry their own set of real-world complications: tracked changes, comments, embedded images, headers and footers, tables, and multiple levels of styling that don’t always map cleanly onto a linear text extraction.

A particular thing worth checking for explicitly: whether tracked changes and comments are being included in extracted text. Depending on the parsing library and its configuration, deleted (but not yet accepted) text can sometimes get extracted alongside the current, correct version — which silently corrupts your knowledge base with outdated or explicitly-rejected content unless you verify your extraction is reading the final, accepted version of the document, not its full edit history.


Parsing CSV and Tabular Data

CSV and other tabular formats present a fundamentally different ingestion challenge than the document formats above: the data isn’t prose to begin with, so “extraction” isn’t really the task — the task is deciding how to meaningfully represent structured rows and columns as retrievable text.

Naively dumping a CSV’s raw rows as text ("John, 34, Engineering, 2019") loses the column context that gives those values meaning. A better approach converts each row into a natural-language sentence using its column headers ("John is 34 years old, works in Engineering, and joined in 2019"), which is far more embeddable and semantically meaningful for retrieval — a plain embedding model has a much easier time relating a sentence to a query than relating an unlabeled sequence of comma-separated values. For genuinely large or highly structured tabular datasets, it’s also worth asking whether the data belongs in a vector database’s similarity search at all, versus a traditional structured query — a question we’ll pick up again when we cover query routing later in this series.


Ingesting from Databases and APIs

Not all source data starts as a file — a significant share of real-world RAG knowledge bases pull directly from live systems: relational databases, internal APIs, SaaS platforms like a CRM or ticketing system.

The core ingestion challenge here is less about text extraction (the data is usually already structured and machine-readable) and more about synchronization: source systems change continuously, and a RAG knowledge base built from a one-time export goes stale the moment the underlying data changes. This connects directly back to the embedding drift and CRUD discussion from Vector DB — a production ingestion pipeline pulling from databases or APIs needs an ongoing sync strategy (scheduled polling, or ideally event-driven updates triggered by the source system itself) rather than a single manual export, or your RAG system will confidently serve increasingly outdated answers without any visible error.

A second, easily overlooked concern with API and database ingestion specifically: access control. If a database or API enforces per-user or per-role permissions, that access boundary needs to be preserved somewhere in your ingestion and retrieval pipeline — otherwise your RAG system risks surfacing information to a user that they were never supposed to see in the source system, silently defeating whatever permissions model the original data was protected by.


Data Cleaning

Regardless of source format, extracted text almost always needs a cleaning pass before it’s ready to embed — and skipping this step is one of the most common, avoidable causes of poor retrieval quality.

Common cleaning tasks worth doing deliberately rather than assuming raw extraction is good enough: removing boilerplate (headers, footers, page numbers, repeated navigation text that leaked through extraction), normalizing whitespace and encoding issues (extracted PDF and OCR text is especially prone to broken spacing, stray line breaks mid-sentence, and inconsistent character encoding), de-duplicating near-identical content (common when the same document exists in multiple versions or locations across your source systems), and stripping content that’s structurally present but semantically empty (table-of-contents entries, disclaimer boilerplate repeated on every page).

The broader principle worth carrying forward: cleaning isn’t a mechanical formality — it’s a genuine quality lever, arguably as consequential to final retrieval quality as the choice of embedding model . A well-chosen embedding model applied to messy, boilerplate-laden text will consistently underperform a modest embedding model applied to clean, well-structured text.


Closing Thoughts

Ingestion is where a RAG system’s ceiling actually gets set — every downstream post in this series (chunking, embedding, retrieval, generation) can only work with the text ingestion actually managed to extract cleanly. Treating this stage as a quick, disposable preprocessing step rather than a genuine engineering problem is one of the most common reasons ambitious RAG systems underperform in production despite using strong models everywhere else.

With clean, extracted text now in hand, next post picks up the very next question: how do you split that text into the right-sized, well-bounded chunks that actually get embedded and retrieved — and why getting chunk size and boundaries wrong can sabotage even a perfectly clean ingestion pipeline.