Hybrid Search: BM25 + Vectors with Reciprocal Rank Fusion

October 10, 2026 · 4 min read

Search developer docs for "fetch ECONNRESET retry". The best page is titled "ECONNRESET error codes" — it contains the exact rare token you typed. A vector search can easily miss it: to an embedding model, ECONNRESET is a few unfamiliar sub-word pieces with little meaning. It will find "Handling flaky networks", which never mentions the token but is about the same thing.

Keyword search has the opposite blind spot. Search for "my request keeps dying halfway" and it finds nothing useful — none of those words appear on the pages about connection resets.

Hybrid search runs both and merges the results. In RAG pipelines, it's one of the cheapest reliable improvements to retrieval.

Keyword side: BM25

BM25 is the standard keyword ranking function — the default in Elasticsearch, OpenSearch and Lucene. For each query term in a document, it combines three ideas:

  • Rare terms matter more (inverse document frequency). "ECONNRESET" appearing in 3 of 10,000 documents is a strong signal; "fetch" in 2,000 is weak.
  • Repetition helps, with diminishing returns. The tenth occurrence of a term adds much less than the second. A parameter k1 (usually 1.2–2.0) controls how fast the score saturates.
  • Long documents are penalised a little, because they contain more words by chance. A parameter b (usually 0.75) controls how much.

It needs no model, no GPU and no embeddings. It handles exact identifiers, error codes, product SKUs and names very well, and it explains itself: a document scores because it contains these terms.

Semantic side: vectors

Embeddings map text to vectors so that similar meanings land close together. Vector search returns the nearest neighbours by cosine similarity. It handles synonyms, paraphrases and questions phrased nothing like the answer — and struggles with rare tokens, exact numbers and negation.

Merging: why you can't add scores

The obvious merge is bm25Score + cosineSimilarity. It doesn't work: BM25 scores are unbounded and vary by query (8.4 for one, 31 for the next); cosine similarities sit in a narrow band, often 0.7–0.9. Adding them lets BM25 decide almost everything. You can normalise each list to 0–1 first, but min-max normalisation is sensitive to outliers and to how many results you fetched.

Reciprocal rank fusion (RRF) sidesteps scales entirely by using only positions. Each list gives a document 1 / (k + rank), and the scores are summed:

function rrf(lists: string[][], k = 60) {
const score = new Map<string, number>()
for (const list of lists) {
list.forEach((id, i) => {
const rank = i + 1
score.set(id, (score.get(id) ?? 0) + 1 / (k + rank))
})
}
return [...score].sort((a, b) => b[1] - a[1])
}
 
rrf([bm25(query), vectorSearch(query)])
BM25 (keywords)
1D3 ECONNRESET error codes
2
D1 Retrying failed fetch calls
3D5 fetch() options reference
4D2 Handling flaky networks

Query: "fetch ECONNRESET retry". BM25 scores exact terms, so the page that literally contains the rare token ECONNRESET ranks first.

0 / 3

The constant k (60 in the original paper, by Cormack, Clarke and Büttcher) damps the advantage of the very top ranks. With k = 60, rank 1 is worth 1/61 and rank 10 is worth 1/70 — not very different. So a document that's fairly high in both lists beats one that's first in one list and absent from the other. That's what you want from two imperfect signals.

function rrf(lists: string[][], k = 60): [string, number][] {
  const score = new Map<string, number>()
  for (const list of lists) {
    list.forEach((id, i) => {
      score.set(id, (score.get(id) ?? 0) + 1 / (k + i + 1))
    })
  }
  return [...score].sort((a, b) => b[1] - a[1])
}

const results = rrf([await bm25(query, 50), await vectorSearch(query, 50)])

RRF has no training, no tuning beyond k, and works for any number of lists — add a third retriever (titles only, or a different embedding model) by adding a list. Many search engines and vector databases ship it built in.

If you have labelled queries, a weighted combination of normalised scores can beat RRF. Measure it with an eval set before switching; without one, RRF is the safer default.

Then rerank

Fusion decides which ~50 candidates make the shortlist. A reranker — a cross-encoder that reads the query and each document together — then orders the shortlist much more accurately than either retriever. It's too slow to run over the whole corpus, so it's used only on the shortlist.

query ─┬─▶ BM25 top 50 ───┐
       └─▶ vector top 50 ─┴─▶ RRF ─▶ top 50 ─▶ reranker ─▶ top 5 ─▶ LLM

Practical notes

  • Fetch more than you need from each side (50–100), so a document ranked 30th by one retriever can still be rescued by the other.
  • Keep the tokenisers sensible. BM25 is only as good as its analyser: splitting ECONNRESET or user_id into pieces, stemming code identifiers, or dropping numbers as "stop words" all hurt technical search.
  • Look at failures by type. Queries with identifiers, error codes and names usually need BM25; vague natural-language questions need vectors. Hybrid search is what lets one search box serve both.