Natural Language Processing

Vector Databases Explained: Where Your Embeddings Actually Live

What a vector database really is under the hood — storage architecture, indexing, metadata filtering, and the operational details that separate a demo from a production RAG system.

Vector Databases

Previous post gave us the pieces: text turned into vectors, and models that do the turning. But vectors sitting in a Python list or a Pandas dataframe don’t scale past a toy demo. The moment your collection grows past a few thousand documents, you need something purpose-built to store these vectors, search them fast, and keep them organized alongside the rest of your data. That something is a vector database — and this post is about what it’s actually doing under the hood, not just how to call its API.


Why Vector Databases Exist

Here’s the naive approach: store every embedding in a list, and when a query comes in, compute the similarity between the query vector and every single stored vector, then sort and return the top matches. This is called a brute-force or exact nearest neighbor search, and it works completely fine — for a few thousand vectors.

The problem is that this approach scales linearly: double your document collection, double your search time. At a hundred thousand documents it’s noticeably slow. At tens of millions, it’s unusable for anything resembling a responsive application. A traditional relational database doesn’t help either — B-tree indexes, the workhorse behind fast SQL lookups, are built for exact matches and range queries on scalar values, not for finding “nearest” points in a 1,000-dimensional space.

Vector databases exist to solve exactly this gap: they use specialized approximate nearest neighbor (ANN) indexing structures (which we’ll unpack below) to return results that are almost certainly the true closest matches, in a fraction of the time brute-force search would take — trading a small, tunable amount of accuracy for a massive gain in speed. That trade-off is the entire reason this category of database exists.


Storage Architecture

It helps to think of a vector database as holding two parallel layers of information for every entry:

  • The vector itself — the raw embedding, typically hundreds to thousands of floating-point numbers, usually stored in a format optimized for fast bulk reads (rather than a general-purpose row format).
  • The associated payload — the original text (or a reference to it), plus any metadata (which we’ll cover shortly), stored alongside the vector so a search result can return something a human — or an LLM — can actually use, not just a list of numbers.

Underneath this, most vector databases maintain the ANN index (the structure that makes fast search possible) as a separate, purpose-built layer from the raw vector storage — the index doesn’t contain the vectors themselves so much as a navigable map of how they relate to each other in space, built specifically to be traversed quickly during a query.

This two-layer design — raw storage plus a separate search-optimized index — is why vector databases are a distinct category of infrastructure rather than just “a table with a vector column bolted on,” even though some general-purpose databases (like PostgreSQL with the pgvector extension) now offer that as an added capability.


Collections

A collection (sometimes called an “index” or “table,” depending on the specific database) is a named, logically separate group of vectors — the vector-database equivalent of a table in a relational database.

Collections matter for two practical reasons. First, isolation: you typically want your product documentation embeddings in a different collection from your customer support tickets, since mixing unrelated content in one search space means a query about a return policy might surface a paragraph from a technical manual purely because it happened to be geometrically nearby. Second, configuration: each collection can specify its own vector dimensionality, distance metric (cosine, dot product, Euclidean, etc), and index settings — because different embedding models and use cases genuinely need different configurations.

A useful rule of thumb for RAG systems: one collection per genuinely distinct knowledge domain or embedding model, not one collection per customer or per document — that granularity is usually better handled by the metadata and namespace mechanisms covered next.


Metadata

Metadata is structured, non-vector information attached to each stored entry — things like a source filename, publication date, author, document type, access permissions, or a chunk’s position within its parent document.

This matters enormously in practice because pure vector similarity alone often isn’t a sufficient answer to a real query. Consider “What was our refund policy last quarter?” — semantic similarity alone can find passages about refund policy, but has no native way to also constrain results to a specific time period. Metadata bridges this gap, letting you combine semantic search with the kind of precise, structured constraints a traditional database query handles naturally.

Well-designed metadata is one of the most underrated levers in RAG retrieval quality — it’s often cheaper and more reliable to fix a retrieval problem by adding a useful metadata field than by switching to a fancier embedding model.


Indexing

This is the mechanical core of what makes a vector database fast: the ANN index structure referenced above. The dominant approach in modern vector databases is HNSW (Hierarchical Navigable Small World) — worth understanding at least at a conceptual level, since it’s the default in most systems you’ll actually use.

HNSW builds a multi-layered graph where each vector is a node, connected to other vectors that are close to it in the embedding space. The top layer is sparse, with long-range connections spanning large distances across the space; each layer below gets progressively denser with shorter-range connections. A search starts at the sparse top layer, quickly narrows down to the right general neighborhood, then descends layer by layer, refining the search until it lands on the nearest actual neighbors — conceptually similar to how you’d navigate a country by first picking the right region on a map, then the right city, then the right street, rather than scanning every address in the country individually.

Other indexing approaches exist — IVF (Inverted File Index, which clusters vectors into buckets and searches only the most promising buckets) and product quantization (which compresses vectors to save memory, at some cost to precision) are common enough to recognize by name, sometimes used in combination with HNSW rather than as a replacement for it.

The practical trade-off across all of these: better index quality (more accurate results) generally costs more memory and slower index-build time. Most vector databases expose tunable parameters here — worth understanding that “accuracy” in an ANN index is a dial you can turn, not a fixed guarantee.


Filtering

Filtering is the operational mechanism that makes metadata actually useful at query time: restricting a similarity search to only the subset of vectors matching specific metadata conditions — for example, “find the most semantically similar passages, but only among documents tagged department: legal and date > 2025-01-01.”

There’s a real engineering subtlety here worth knowing about, because it affects both speed and accuracy: pre-filtering narrows the candidate set by metadata before running the similarity search, while post-filtering runs the similarity search first and discards non-matching results afterward. Post-filtering is simpler to implement but can be a serious problem in practice — if your filter is narrow (say, matching only 1% of documents) but you only retrieved the top 10 similarity results before filtering, you might end up with zero results even though good matches exist elsewhere in the collection. Good vector databases implement filtering that’s aware of this and integrates the filter directly into the graph traversal, rather than bolting it on as an afterthought.


Namespaces

A namespace is a lighter-weight partition than a full collection — a way of segmenting vectors within the same collection, typically used for per-tenant or per-user isolation in multi-tenant applications.

The distinction from collections matters operationally: collections usually involve separate configuration (dimensionality, distance metric, index settings) and are the right tool for genuinely different data types. Namespaces share all of that configuration but keep data logically separated — the right tool when you have structurally identical data (say, every customer’s uploaded documents follow the same schema and embedding model) that simply needs to stay isolated from other customers’ data, without spinning up a new collection per customer.

If you’re building a RAG product that serves multiple customers or users from shared infrastructure, namespaces are usually the mechanism that keeps one customer’s query from ever seeing another customer’s private documents — worth verifying explicitly with whichever vector database you choose, since the isolation guarantee is a security property, not just an organizational convenience.


CRUD Operations

Like any database, a vector database supports the standard Create, Read, Update, Delete lifecycle — but each takes on a slightly different character than in a traditional relational database, worth knowing before you assume the two behave identically.

  • Create (upsert) — inserting a new vector plus its metadata. Most vector databases expose this as an “upsert” (update-or-insert) rather than a strict insert, since re-embedding and re-inserting an updated document is a far more common pattern than a true first-time insert.
  • Read (query) — the similarity search itself: given a query vector (and optional metadata filters), return the top-k nearest results.
  • Update — modifying metadata or replacing a vector. Note that updating the text behind a vector generally requires re-embedding and re-inserting the whole entry — you can’t partially edit a vector’s meaning the way you’d edit a text field, since the vector is a single computed representation of the whole content.
  • Delete — removing a vector and its metadata. This directly connects back to the embedding drift problem from Embeddings’s post, whenever a source document changes or is removed, its old vector needs to be deleted too, or stale, outdated content keeps getting retrieved and handed to your LLM indefinitely — a silent failure mode that’s easy to overlook.

Getting CRUD operations right is less glamorous than picking the right index or embedding model, but it’s where a surprising number of production RAG systems quietly fail: content gets updated in the source system but never re-synced to the vector database, and the LLM keeps confidently answering from data that's no longer true.


Closing Thoughts

A vector database isn’t magic — it’s a purpose-built combination of storage, an ANN index for speed, and metadata handling for precision, wrapped in an interface that feels like a familiar database. Understanding what’s actually happening at each of these layers — collections, indexing, filtering, namespaces, and the CRUD lifecycle — is what lets you debug a slow or inaccurate RAG system instead of just tweaking parameters at random.