AI / ML · Retrieval & Embeddings
Embeddings + vector DBs without the magic
What an embedding actually is, how vector search works under the hood, and how to pick between pgvector, Qdrant, and Pinecone.
An embedding is a list of numbers. Usually 384, 768, or 1536 of them. The model that produced it was trained so that semantically-similar inputs end up close to each other in this number-space, and dissimilar inputs end up far apart.
That's it. Everything else is bookkeeping.
How "close" works#
Three distance functions matter:
- Cosine similarity — angle between two vectors, ignores magnitude. Range
[-1, 1]. The default for almost everything. - Dot product — like cosine but sensitive to magnitude. Useful when magnitude carries info (e.g. some IR-style scoring schemes).
- Euclidean (L2) — straight-line distance. Used when vectors are normalized to unit length (then it's mathematically equivalent to cosine).
If your embedding model normalizes outputs (most do — OpenAI, Voyage, Cohere all return unit vectors), cosine, dot product, and L2 give the same ranking. Don't sweat the choice unless you're getting weird results.
What gets trained#
Embedding models are trained on contrastive objectives: given a query and a "positive" (relevant) passage, plus several "negatives" (irrelevant passages), pull the query close to the positive and push it away from the negatives.
The training data shape determines what the model thinks similarity means. A model trained on (question, answer) pairs will score "What's the capital of France?" close to "Paris is the capital." A model trained on (paraphrase pairs) will score the question close to "France's capital — what is it?". Different objectives, different latent geometry, different downstream behavior.
Practical implication: don't mix embedding models. Embeddings from text-embedding-3-small are not comparable to embeddings from voyage-3 even though they're both 1536-dim. Re-embed your whole corpus when you switch.
Picking a model#
The MTEB leaderboard (huggingface.co/spaces/mteb/leaderboard) ranks open and closed models on real benchmarks. As of mid-2026:
| Model | Dim | Cost / 1M tokens | Notes |
|---|---|---|---|
text-embedding-3-small (OpenAI) |
1536 | $0.02 | Strong default. Cheap. |
text-embedding-3-large (OpenAI) |
3072 | $0.13 | Marginal gain over -small. |
voyage-3-large |
1024 | $0.18 | Top retrieval performance. |
cohere-embed-v3 |
1024 | $0.10 | Good multilingual. |
bge-large-en-v1.5 (open) |
1024 | self-host | Best of the open family. |
For most products: start with text-embedding-3-small. It's good, it's cheap, and switching later is a one-time corpus re-embed.
Vector indexes: the speed/recall tradeoff#
A naive scan over N embeddings is O(N) per query. Fine for 10K vectors, fatal for 10M. Vector indexes trade exact recall for speed.
The two index families that matter:
HNSW (Hierarchical Navigable Small World) — a graph where each node is a vector, edges connect "near" vectors. Search walks the graph greedily. Recall is high (95%+) at reasonable speed. Memory-hungry: indexes are 2–4× the raw vector size. The default for most modern vector DBs.
IVF (Inverted File) — clusters vectors into K lists by k-means; query searches the M nearest lists. Faster build, lower memory, lower recall (88–92% typical). Good when you have hundreds of millions of vectors and can afford to lose a bit of recall.
In pgvector both are available:
-- HNSW: better recall, slower to build
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
-- IVFFlat: faster build, less memory
CREATE INDEX ON chunks USING ivfflat (embedding vector_cosine_ops)
WITH (lists = 100);m in HNSW controls graph connectivity; ef_construction controls index quality. Defaults are usually fine.
Vector DB comparison: pgvector vs. Pinecone vs. Qdrant#
For most teams, pgvector is the right answer. You already run Postgres. You already have backups. It scales to ~10M vectors comfortably on a single node. Hybrid queries with BM25 and SQL filters are trivial — same SQL, same transaction.
When pgvector breaks down:
- >50M vectors — index size pushes Postgres memory hard. Migrate to a dedicated store.
- Sub-50ms p99 with high QPS — pgvector contends with regular workload. A dedicated vector DB on its own hardware wins.
- Multi-tenancy with isolation per tenant — pgvector's per-row WHERE filtering works but isn't as fast as Pinecone's namespaces or Qdrant's collections.
- pgvector10
- Qdrant100
- Pinecone500
- LanceDB50
- Weaviate80
| pgvector | Qdrant | Pinecone | LanceDB | |
|---|---|---|---|---|
| Hosted | self | self+cloud | cloud-only | self |
| SQL filter | yes (native) | yes | yes (limited) | yes |
| Hybrid (vector + BM25) | yes | yes | no | no |
| p99 at 10M vectors | ~50ms | ~10ms | ~5ms | ~30ms |
| Cost (10M @ light load) | $30/mo (Postgres) | $80/mo | $250/mo | $0 (file-based) |
| Operational complexity | low (one DB) | medium | low (managed) | low (no server) |
When you don't need a vector DB#
- Corpus < 100K chunks: numpy array +
argsort(np.dot(matrix, query))on every request. 10ms for ~50K. Real shipped systems run on this. - Static corpus that fits in memory: LanceDB or even just a parquet file + DuckDB.
- Read-heavy, write-rarely: precompute top-K nearest neighbors per chunk offline; serve a lookup. No vector search at request time at all.
The fastest "vector DB" is no vector DB.
Pitfalls#
A few things that bite first-timers:
- Different distance functions per index type. pgvector ivfflat doesn't support cosine without normalized vectors; check the docs.
- Drift after model upgrade. When the embedding provider releases a new model and you upgrade, your old embeddings are still compatible with each other, but new queries embedded with the new model can't be compared against them. Re-embed atomically.
- Filters before vector search vs. after. Filtering after retrieval can produce empty results when the filter is selective. Most vector DBs support pre-filtering — use it.
- Index rebuilds block writes. In pgvector,
REINDEXblocks. Plan accordingly. Concurrent rebuilds (REINDEX CONCURRENTLY) work in PG14+.
Further reading#
- sbert.net — the canonical embedding-models reference, comparison, and best practices.
- Pinecone's vector search guide — vendor-y but the technical content is solid.
- pgvector README — small, readable, definitive on the Postgres side.
- "Embedding fine-tuning guide" by Cohere — when stock embeddings underperform on your corpus, fine-tuning is sometimes the answer.