Search developer docs for "fetch ECONNRESET retry". The best page is
titled "ECONNRESET error codes" — it contains the exact rare token you
typed. A vector search can easily miss it: to an embedding model,
ECONNRESET is a few unfamiliar sub-word pieces with little meaning. It
will find "Handling flaky networks", which never mentions the token but is
about the same thing.
Keyword search has the opposite blind spot. Search for "my request keeps dying halfway" and it finds nothing useful — none of those words appear on the pages about connection resets.
Hybrid search runs both and merges the results. In RAG pipelines, it's one of the cheapest reliable improvements to retrieval.
Keyword side: BM25
BM25 is the standard keyword ranking function — the default in Elasticsearch, OpenSearch and Lucene. For each query term in a document, it combines three ideas:
- Rare terms matter more (inverse document frequency). "ECONNRESET" appearing in 3 of 10,000 documents is a strong signal; "fetch" in 2,000 is weak.
- Repetition helps, with diminishing returns. The tenth occurrence of a
term adds much less than the second. A parameter
k1(usually 1.2–2.0) controls how fast the score saturates. - Long documents are penalised a little, because they contain more
words by chance. A parameter
b(usually 0.75) controls how much.
It needs no model, no GPU and no embeddings. It handles exact identifiers, error codes, product SKUs and names very well, and it explains itself: a document scores because it contains these terms.
Semantic side: vectors
Embeddings map text to vectors so that similar meanings land close together. Vector search returns the nearest neighbours by cosine similarity. It handles synonyms, paraphrases and questions phrased nothing like the answer — and struggles with rare tokens, exact numbers and negation.
Merging: why you can't add scores
The obvious merge is bm25Score + cosineSimilarity. It doesn't work: BM25
scores are unbounded and vary by query (8.4 for one, 31 for the next);
cosine similarities sit in a narrow band, often 0.7–0.9. Adding them lets
BM25 decide almost everything. You can normalise each list to 0–1 first,
but min-max normalisation is sensitive to outliers and to how many results
you fetched.
Reciprocal rank fusion (RRF) sidesteps scales entirely by using only
positions. Each list gives a document 1 / (k + rank), and the scores
are summed:
Query: "fetch ECONNRESET retry". BM25 scores exact terms, so the page that literally contains the rare token ECONNRESET ranks first.
The constant k (60 in the original paper, by Cormack, Clarke and Büttcher)
damps the advantage of the very top ranks. With k = 60, rank 1 is worth
1/61 and rank 10 is worth 1/70 — not very different. So a document that's
fairly high in both lists beats one that's first in one list and absent
from the other. That's what you want from two imperfect signals.
function rrf(lists: string[][], k = 60): [string, number][] {
const score = new Map<string, number>()
for (const list of lists) {
list.forEach((id, i) => {
score.set(id, (score.get(id) ?? 0) + 1 / (k + i + 1))
})
}
return [...score].sort((a, b) => b[1] - a[1])
}
const results = rrf([await bm25(query, 50), await vectorSearch(query, 50)])
RRF has no training, no tuning beyond k, and works for any number of
lists — add a third retriever (titles only, or a different embedding model)
by adding a list. Many search engines and vector databases ship it built
in.
If you have labelled queries, a weighted combination of normalised scores can beat RRF. Measure it with an eval set before switching; without one, RRF is the safer default.
Then rerank
Fusion decides which ~50 candidates make the shortlist. A reranker — a cross-encoder that reads the query and each document together — then orders the shortlist much more accurately than either retriever. It's too slow to run over the whole corpus, so it's used only on the shortlist.
query ─┬─▶ BM25 top 50 ───┐
└─▶ vector top 50 ─┴─▶ RRF ─▶ top 50 ─▶ reranker ─▶ top 5 ─▶ LLM
Practical notes
- Fetch more than you need from each side (50–100), so a document ranked 30th by one retriever can still be rescued by the other.
- Keep the tokenisers sensible. BM25 is only as good as its analyser:
splitting
ECONNRESEToruser_idinto pieces, stemming code identifiers, or dropping numbers as "stop words" all hurt technical search. - Look at failures by type. Queries with identifiers, error codes and names usually need BM25; vague natural-language questions need vectors. Hybrid search is what lets one search box serve both.