RAG & Knowledge Systems · 6 min
Hybrid Retrieval: Why Lexical BM25 Still Beats Embeddings Sometimes
BM25 ranks by exact-token statistics, so it nails rare identifiers and error codes that dense vectors blur; here's how to fuse it with embeddings and cover both failure modes.
If you have only ever built retrieval on an embedding model, you have quietly accepted a blind spot. Dense vectors are trained to collapse surface form into meaning: CVE-2021-44228, Log4Shell, and "that Java logging RCE" all land near each other in the vector space. That is exactly what you want when a user paraphrases. It is exactly what you do not want when the user pastes ORA-01555 and needs the one document that contains that literal string. The encoder has already smeared ORA-01555 toward ORA-01552 and every other Oracle error, because at the subword level those tokens look nearly identical and were too rare in training to pull apart.
BM25 has no such problem, because BM25 never learned anything. It is bookkeeping over exact tokens.
The mental model: BM25 counts, it does not understand
BM25 (Robertson & Zaragoza, 2009) scores a document against a query by summing a contribution from each query term the document actually contains. Three quantities drive each term:
- Term frequency
f(q, D): how many times the term appears in the document. More is better, with diminishing returns. - Inverse document frequency
IDF(q): how rare the term is across the corpus. A term in 3 of 10 million docs is gold; a term in 9 million is noise. - Length normalization: a 50-word doc that mentions the term twice is more "about" it than a 5,000-word doc that mentions it twice.
The scoring function ties them together:
f(q,D) · (k1 + 1)
BM25(q,D) = Σ IDF(q) · ───────────────────────────────────
q∈Q f(q,D) + k1 · (1 − b + b · |D|/avgdl)
k1 term-frequency saturation (Lucene default 1.2)
b length-normalization weight (Lucene default 0.75, 0=off)
|D| doc length, avgdl = average doc length in corpus
IDF(q) = log(1 + (N − n(q) + 0.5) / (n(q) + 0.5)) N = #docs, n(q) = #docs containing qTwo parameters shape the curve. k1 controls saturation: with k1 = 1.2, the jump from 1 to 2 occurrences matters a lot, from 20 to 21 almost nothing. A term mentioned once is a hit; mentioned fifty times it is spam, not fifty hits. b controls how hard long documents get penalized. That is the whole model. No training, no GPU, no drift when your domain shifts.
The IDF term is where the magic for rare strings lives. Because IDF grows as a token gets rarer, an identifier that appears in a single document earns an enormous weight, and BM25 ranks that document first almost by construction. A dense retriever would have to have learned that ORA-01555 matters, and it usually hasn't, because the token was too rare to earn a good embedding.
Where each one breaks
Line them up honestly:
Query BM25 Dense embeddings ────────────────────────────────────────────────────────────── "ORA-01555 snapshot too old" nails it blurs to nearby ORA-* exact function name foo_bar() exact match may miss, tokenized oddly "how do I roll back a txn" weak (few strong (understands → doc says "revert a shared tokens) rollback≈revert) transaction" non-English / typo'd query brittle more robust long tail rare vocabulary strong weak (undertrained tokens)
BM25's failure mode is vocabulary mismatch. If the query and the document use different words for the same idea, the sum runs over an empty set of shared terms and the score falls near zero. "Revert a transaction" and "roll back a txn" share almost no tokens. Dense retrieval eats that case for breakfast.
Dense retrieval's failure mode is the mirror image: exact, rare, out-of-distribution tokens. This is not a small edge case. The BEIR benchmark (Thakur et al., 2021) evaluated retrieval systems zero-shot across 18 datasets and found BM25 a shockingly hard baseline to beat out-of-domain. Dense models that crushed BM25 on their training distribution fell behind it on unfamiliar corpora. The community shorthand became "BM25 is hard to beat zero-shot."
Defender: Your log lines, stack traces, config keys, and CVE IDs are exactly the rare high-entropy tokens embeddings handle worst. A pure-vector RAG over an incident-response knowledge base will confidently miss the doc that names the exact error. Keep a lexical index in the loop or you retrieve vibes instead of the runbook.
Fusing the two
You want a recall set that covers both failure modes, so run both retrievers and merge. The naive merge, adding the BM25 score to the cosine similarity, does not work: the two scores live on incompatible scales. BM25 is unbounded and corpus-dependent (scores of 8, 19, 40 and up); cosine sits in [-1, 1]. Summing them lets whichever number happens to be bigger dominate for no principled reason.
Two ways out.
Score normalization. Min-max each retriever's scores into [0, 1] over the candidate set, then take a weighted sum: α · norm(dense) + (1 − α) · norm(bm25). It works, but it is fragile. One outlier stretches the min-max range and squashes everything else, and the range shifts every query, so α never quite transfers.
Reciprocal Rank Fusion (RRF) (Cormack, Clarke & Büttcher, 2009) throws the scores away and fuses ranks:
RRF(d) = Σ 1 / (k + rank_r(d)) k = 60 (the value from the paper)
r∈R
R = the retrievers (BM25, dense, …)
rank_r(d) = position of doc d in retriever r's list (1 = top), if presentRank is unitless, so the scale problem evaporates. The k = 60 constant is a smoother: it flattens the gap between rank 1 and rank 2 so a single retriever's top hit cannot steamroll the fusion, and it rewards documents that multiple lists agree on. A doc at rank 3 in both lists beats a doc at rank 1 in one list and absent from the other. Cormack found 60 empirically on TREC data; anything in roughly 40 to 80 behaves about the same, so most engines just ship 60.
Worked micro-example, two retrievers, k = 60:
doc BM25 rank dense rank RRF score ──────────────────────────────────────────────────────── A 1 30 1/61 + 1/90 = 0.0164+0.0111 = 0.0275 B 2 2 1/62 + 1/62 = 0.0161+0.0161 = 0.0323 ← wins C — 1 0 + 1/61 = 0.0164
Doc B wins despite topping neither list, because both retrievers vouch for it. Doc C is dense's darling, but BM25 never saw it, so it ranks third. That consensus behavior is the whole point. The fused set carries the exact-match hits and the paraphrase hits, and the reranker (or the LLM) sorts out the rest.
Builder: RRF is about ten lines of code with no tuning, no score calibration, no per-query weights. Start here. Reach for weighted score normalization only once you have measured that you need to lean harder on one retriever, and you probably don't.
Researcher: RRF is order-only. It discards the margin between ranks 1 and 2, which sometimes carries real signal. Weighted score fusion with per-retriever normalization, or a small learned fusion layer, can beat it when your two retrievers have very different reliability. Measure on your own queries. The BEIR lesson is that retrieval quality is brutally domain-dependent and nobody's leaderboard number is your number.
The practical build: index your corpus twice, an inverted index (any Lucene-family engine, or Postgres FTS at small scale) and a vector index. Retrieve the top ~50 from each, fuse with RRF, then rerank the fused ~100 down to the handful you feed the model. This is the standard chunking-to-context pipeline from the earlier retrieval lesson, with one change that costs a few milliseconds and buys back an entire class of misses. "Just embed everything" is the default that quietly loses your error codes.
Sources
- Robertson, S. & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4). https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf
- Cormack, G., Clarke, C. & Büttcher, S. (2009). Reciprocal Rank Fusion outperforms Condorcet and individual Rank Learning Methods. SIGIR 2009. https://dl.acm.org/doi/10.1145/1571941.1572114
- Thakur, N., Reimers, N., Rücklé, A., Srivastava, A. & Gurevych, I. (2021). BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets & Benchmarks. https://arxiv.org/abs/2104.08663
- Apache Lucene,
BM25Similarity(defaultk1 = 1.2,b = 0.75;IDF = log(1 + (N − n + 0.5)/(n + 0.5))). https://lucene.apache.org/core/9_0_0/core/org/apache/lucene/search/similarities/BM25Similarity.html