Natural Language Processing

Embedding Models: From Word2Vec to Voyage AI — A Practical Field Guide

A grounded tour of the embedding models that actually matter — how we got from static word vectors to today's Voyage, OpenAI, and BGE-class models, and a framework for choosing between them.

Embedding Models

Developer Diaries: Building with Retrieval-Augmented Generation

Previous post covered what embeddings are — vectors that place similar meaning close together in space. This chapter is about who actually builds those vectors: the specific models behind every “embed this text” API call you’ll make while building a RAG system, how they evolved, and — the question that actually matters when you’re shipping something — which one to reach for.


The Static Era: Word2Vec, GloVe, and FastText

Modern embeddings trace back to a cluster of ideas from the early 2010s that first made the “words with similar meaning get similar vectors” property practical at scale.

Word2Vec (2013) was the breakthrough that popularized the whole idea. It learned word vectors by training a shallow neural network to predict a word from its surrounding context (or vice versa) across a massive text corpus — and the resulting vectors captured relationships that felt almost magical at the time: the vector arithmetic king - man + woman landed close to queen, purely as an emergent property of co-occurrence statistics.

GloVe (2014), developed at Stanford, took a different mathematical route to a similar destination — instead of predicting context word-by-word, it factorized a global word co-occurrence matrix directly. In practice, GloVe and Word2Vec produce embeddings of comparable quality; the interesting part is that two different training objectives converged on the same underlying idea.

FastText (2016), from Facebook AI, addressed a specific weakness in both: neither Word2Vec nor GloVe could produce a sensible vector for a word they hadn’t seen during training. FastText fixed this by representing words as bags of character n-grams rather than atomic units — so it could construct a reasonable embedding for misspellings, rare words, or entirely unseen words by composing them from familiar sub-word pieces.

All three share the same fundamental limitation flagged in Chapter 4: one fixed vector per word, regardless of context. “Bank” gets the same embedding whether it means a riverbank or a financial institution. This is precisely the ceiling that motivated the next leap.


The Contextual Shift: BERT

BERT (Bidirectional Encoder Representations from Transformers, 2018, Google) is the model that broke the one-word-one-vector ceiling. Built on the transformer architecture, BERT generates a different embedding for a word depending on the full sentence around it — solving exactly the ambiguity problem static embeddings couldn’t.

BERT was trained on two clever self-supervised tasks that needed no manual labeling: predicting randomly masked-out words from context (masked language modeling), and predicting whether one sentence plausibly follows another (next sentence prediction). This let it learn deep, bidirectional language understanding from raw text alone, at a scale no labeled dataset could match.

There’s a catch worth understanding, though: BERT wasn’t actually designed to produce standalone embeddings for retrieval. Its native output is a set of per-token vectors optimized for tasks like classification or question answering, not a single well-behaved vector representing an entire sentence’s meaning. Naively averaging BERT’s token vectors to get a sentence embedding produces mediocre results — which set up the exact problem the next model was built to solve.


Making Sentence Embeddings Practical: SBERT

SBERT (Sentence-BERT, 2019) modified BERT specifically to produce high-quality sentence-level embeddings efficiently — the missing piece for using BERT-style models in semantic search and retrieval.

The core idea: fine-tune BERT using a siamese network structure — running pairs of sentences through the same shared model and training it so that sentences with similar meaning end up with vectors that are close together (measured via cosine similarity), while dissimilar sentences end up far apart. This directly optimizes for the property retrieval actually needs, rather than hoping it emerges as a side effect.

The practical impact was enormous: SBERT made comparing thousands or millions of sentences computationally feasible (you embed once, then compare vectors with cheap similarity math) instead of requiring an expensive full BERT forward pass for every single pair being compared. SBERT is less a single model than a training methodology — and it directly laid the groundwork for essentially every modern sentence/passage embedding model that followed, including all the ones below.


The Modern Open-Model Landscape: E5, BGE, and GTE

By the early 2020s, a cluster of open, high-performing embedding model families emerged — trained on much larger and more carefully curated datasets than SBERT’s original work, and consistently competitive with, or ahead of, commercial APIs on public benchmarks.

E5 (EmbEddings from bidirEctional Encoder rEpresentations, Microsoft) introduced a training approach centered on massive-scale weakly supervised contrastive pretraining, followed by fine-tuning on high-quality labeled pairs. A distinctive detail worth knowing if you use it: E5 models expect text to be explicitly prefixed with "query: " or "passage: " depending on its role — a concrete, practical example of the asymmetric query/passage handling flagged in Embedding’s post.

BGE (BAAI General Embedding, from the Beijing Academy of Artificial Intelligence) is widely used in production for its consistently strong retrieval benchmark performance and permissive open licensing, making it a common default for teams self-hosting embedding infrastructure rather than depending on a commercial API.

GTE (General Text Embeddings, Alibaba) is trained with a multi-stage contrastive learning pipeline and is notable for handling both short queries and long documents well within the same model — a useful property for RAG pipelines that mix short user questions against long-form source content.

The common thread across all three: they’re open-weight, can be self-hosted, and routinely land at or near the top of the MTEB leaderboard introduced in Embedding’s post — which is exactly why they show up so often as the default choice in open-source RAG frameworks.


Jina Embeddings

Jina Embeddings (Jina AI) is worth calling out on its own for a specific technical strength: long-context embedding. Most embedding models are trained and evaluated on relatively short text (a sentence or short paragraph); Jina’s models are explicitly built to embed much longer passages — thousands of tokens — while preserving retrieval quality, which matters directly for the passage-chunking trade-offs.

Jina also offers multilingual and multimodal (text-and-image) embedding variants, positioning it well for RAG systems that need to search across mixed languages or content types rather than English-only text.


Commercial Embedding APIs: Voyage AI and OpenAI

Not every team wants to self-host an embedding model — and for many, a hosted API is the simpler, faster path to production. Two names dominate this space.

Voyage AI builds embedding models exclusively — no chat models, no distractions — and is the embedding provider Anthropic officially recommends for Claude-based RAG pipelines. Its voyage-3-large and related models routinely top retrieval-focused benchmarks, and Voyage was an early mover on Matryoshka representation learning — training embeddings so they can be truncated to smaller dimensions (say, 2048 down to 256) with graceful, predictable quality loss, letting you trade storage and speed against quality without switching models. Voyage also ships domain-specialized models (voyage-code, voyage-law, voyage-finance) fine-tuned for technical, legal, and financial retrieval respectively. Worth noting for accuracy: Voyage AI was acquired by MongoDB in February 2025 and now operates as Voyage AI by MongoDB — the “Anthropic-recommended” relationship is a technology partnership, not common ownership.

OpenAI Embeddings (text-embedding-3-small and text-embedding-3-large) remain among the most widely deployed embedding models in production, largely on the strength of ecosystem simplicity — if you’re already calling the OpenAI API, adding embeddings requires no new infrastructure or vendor relationship. Both models also support Matryoshka-style dimension truncation. text-embedding-3-small is the common default for cost-sensitive, high-volume indexing; text-embedding-3-large trades higher cost for meaningfully better retrieval quality, particularly on multilingual text.

The general pattern across both: commercial APIs optimize for “good enough quality with zero infrastructure,” while open models optimize for “best possible quality-per-dollar if you’re willing to host it yourself.” Neither is universally correct — it depends on your team’s constraints, which brings us to the actual decision framework.


Choosing an Embedding Model

With this many credible options, picking one comes down to a small number of concrete questions rather than chasing whichever model tops the leaderboard this month — MTEB rankings shift constantly, and the gap between the top handful of models is usually small enough that other factors dominate in practice.

  • Self-hosted or API? Self-hosting (BGE, E5, GTE, Jina) gives you full data control and no per-token cost at scale, at the price of managing GPU infrastructure yourself. APIs (Voyage, OpenAI) trade that operational burden for simplicity and usage-based pricing.
  • General-purpose or domain-specific? A model fine-tuned for your domain (code, legal, medical, finance) will typically outperform a strong general model on that domain’s vocabulary and phrasing — worth checking whether a specialized option exists before defaulting to a general-purpose model.
  • Data privacy and residency requirements. If your documents can’t leave your infrastructure — common in regulated industries — self-hosted open models are frequently the only viable option, regardless of how well a commercial API benchmarks.
  • Context length needs. If you’re embedding long passages rather than short chunks, models with longer native context windows (Jina’s long-context variants, for instance) avoid awkward truncation or splitting.
  • Multilingual requirements. Not every embedding model performs evenly across languages — verify multilingual benchmark performance specifically (MIRACL is the standard reference here) rather than assuming general English quality transfers.
  • Cost at your actual scale. Run the math on your expected document volume and re-indexing frequency before committing — the cost difference between a small and large variant of the same model family compounds fast at millions of documents, and matters far less at a few thousand.

The reliable process, echoing point on evaluation: don’t trust a benchmark leaderboard to make this decision for you. Build a small evaluation set from your own representative queries and documents, test your top two or three candidate models against it, and measure retrieval quality (recall@k, nDCG) directly on data that looks like what you’ll actually be searching in production.


Closing Thoughts

The embedding model landscape looks crowded from the outside, but it collapses into a clear lineage once you see the throughline: static word vectors (Word2Vec, GloVe, FastText) gave way to contextual representations (BERT), which needed a dedicated adaptation to produce usable sentence embeddings (SBERT) — and everything since has been scaling and refining that same recipe, whether as an open model you host yourself or an API you call.

With embeddings and embedding models now covered, the next question is mechanical but critical: how do you actually get your raw documents into embeddable, retrievable chunks in the first place? That’s where we’re headed next — chunking strategy, and why getting it wrong quietly sabotages even the best embedding model you could choose.