myibrahim.cloud

AI / ML · Retrieval & Embeddings

RAG patterns that actually work

Naive top-k similarity is where most RAG systems start and where most of them fail. Here's what to do instead.

The first version of every RAG system I've shipped looked the same: split docs into 500-token chunks, embed them with text-embedding-3-small, store in a vector DB, retrieve top-5 by cosine similarity, stuff into the prompt, send to GPT-4. It worked maybe 60% of the time.

The second version of every RAG system I've shipped did something different. Here's what.

Why naive RAG fails#

Vector similarity is a lossy proxy for semantic relevance. Two passages about completely different things can have high cosine similarity if they share enough vocabulary. Two passages about the same thing can have low similarity if they use different words.

Specific failure modes I've watched in production:

  • Query: "How do I cancel my subscription?" → retrieved: a passage about creating subscriptions, because both contain "subscription" tokens densely.
  • Query: "What's our refund policy for B2B contracts?" → retrieved: the consumer refund policy, because the embedding model didn't differentiate.
  • Query: "Show me the migration that added the new index on user_id." → retrieved: random migration files mentioning user_id, none of which were the right one.

In each case the chunk you actually want exists in the corpus. The retriever just couldn't find it.

Hybrid search: BM25 + vectors#

The single biggest upgrade is combining keyword search (BM25) with vector search and merging the results. BM25 is great at exact-term matches; vectors are great at semantic similarity. Each catches what the other misses.

Postgres has both built-in:

-- BM25-ish via tsvector
CREATE INDEX idx_chunks_tsv ON chunks USING gin(to_tsvector('english', content));

-- Vectors via pgvector
CREATE INDEX idx_chunks_emb ON chunks USING ivfflat (embedding vector_cosine_ops);

-- Hybrid retrieval: union of both, scored by reciprocal-rank fusion
WITH bm25 AS (
  SELECT id, ts_rank(to_tsvector('english', content), query) AS score
  FROM chunks, plainto_tsquery('english', $1) AS query
  ORDER BY score DESC LIMIT 50
),
ann AS (
  SELECT id, 1 - (embedding <=> $2::vector) AS score
  FROM chunks ORDER BY embedding <=> $2::vector LIMIT 50
)
SELECT id, SUM(1.0 / (60 + rank)) AS rrf_score
FROM (
  SELECT id, ROW_NUMBER() OVER () AS rank FROM bm25
  UNION ALL
  SELECT id, ROW_NUMBER() OVER () AS rank FROM ann
) AS combined
GROUP BY id ORDER BY rrf_score DESC LIMIT 10;

Reciprocal-rank fusion (RRF) is the under-the-radar tool that makes hybrid easy: each retriever ranks, you sum 1/(k+rank) across retrievers, sort by sum. No score-normalization headache.

Recall@10 across retrieval strategies (internal benchmark)
  • Vectors only58
  • BM25 only49
  • Hybrid (RRF)72
  • Hybrid + Rerank84

Rerank the top-N with a cross-encoder#

After hybrid retrieval gives you 50 candidates, run them through a cross-encoder reranker. This is a smaller model (Cohere Rerank, Voyage Rerank, BGE Reranker) that takes (query, passage) pairs and produces a relevance score. It's slow per-pair but you only run it on 50 candidates, not the whole corpus.

Reranking moved Recall@10 from 72% to 84% on my last project. It costs ~$0.002 per query at Cohere's pricing. Worth it.

from cohere import Client
co = Client()

results = co.rerank(
    query=user_query,
    documents=[chunk.text for chunk in candidates],
    top_n=5,
    model="rerank-english-v3.0",
)
top_chunks = [candidates[r.index] for r in results.results]

Query rewriting: the model is your retriever's friend#

User queries are awful retrieval inputs. They're terse, ambiguous, and context-dependent. Before retrieving, run the query through a small LLM to rewrite it.

Two patterns I use:

HyDE (Hypothetical Document Embeddings): ask the LLM to generate a hypothetical answer to the query, then embed that and retrieve against it. Works because answers and other answers are closer in embedding space than questions and answers.

hypothetical = llm.complete(
    f"Write a passage that would answer this question: {user_query}",
    max_tokens=200,
)
embedding = embed(hypothetical)
chunks = vector_search(embedding)

Query expansion: ask the LLM to generate 3-5 alternative phrasings of the query, embed all of them, retrieve and union. Catches the cases where the user used different vocabulary than the corpus.

Chunking: bigger than you think, with overlap#

The default-everywhere chunking advice — 500 tokens with 50 overlap — is wrong for most domains.

  • Code repos: chunk by function/class, not by token count. A function half-cut between two chunks is useless.
  • Long docs: 1500-token chunks with 200 overlap. Retrieved passages need enough context to be self-contained.
  • Markdown / structured docs: chunk on heading boundaries (##, ###). Each section is a unit.
  • Tables: never chunk a table mid-row. Either the whole table fits in one chunk, or you serialize it as text + retrieve all rows together.

The mental model: a retrieved chunk should be self-contained enough that an LLM looking only at it could answer the query. If the answer requires "knowing what was said three paragraphs earlier," the chunk is too small.

Skip RAG when the context fits#

If your entire corpus fits in the model's context window — and it often does for product docs, internal wikis, or single-codebase Q&A — just put it all in the prompt. Use prompt caching to make it cheap.

A 80K-token corpus, cached, costs ~$0.10 per query at Anthropic's prompt-caching rates. RAG infrastructure costs more in engineering time than that. You can scale to RAG when the corpus crosses ~150K tokens.

The chart below is from a side-by-side I ran last quarter:

Answer accuracy on 100-question eval, %
  • RAG (top-5)65
  • RAG (top-10 + rerank)84
  • Full context (cached)91

Full context wins until your corpus exceeds the window. Then you go to RAG, but you go knowing what you gave up.

Eval, eval, eval#

You cannot tune RAG without an eval set. Build one before tuning anything:

  1. Collect 50–100 real user queries (or imagined queries from your team).
  2. Manually identify which chunks (or which doc passages) should be retrieved for each.
  3. Score each retrieval strategy on Recall@k (did the right chunk make the top-k?) and final-answer correctness (did the LLM produce a correct answer given the retrieved chunks?).
  4. Track these metrics on every change. RAG tuning is iterative; without numbers you're guessing.

Hamel Husain's LLM eval guide is the canonical reference.

Putting it together#

The RAG stack I default to today:

  1. Chunking: domain-aware (code: by function; markdown: by heading; prose: 1500-token + overlap).
  2. Index: BM25 + vector (pgvector for small/medium, Qdrant or LanceDB for >10M chunks).
  3. Retrieve: hybrid via RRF, top-50.
  4. Query rewriting when queries are short (<6 words): HyDE or expansion.
  5. Rerank: top-50 → top-5 via Cohere/BGE.
  6. Generate: stuff top-5 into the prompt with provenance citations (always cite which chunk an answer came from).
  7. Eval: run the full pipeline on a held-out set on every change.

That's the version that hits 80%+ accuracy on real product Q&A. Naive top-k gets you to 60%. The 20-point delta is mostly hybrid + rerank.

Further reading#

  • ai
  • rag
  • retrieval
  • vector-search
  • llms
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.