Hybrid BM25 and vector search ranking methods
Hybrid ranking merges exact keyword matches from BM25 with semantic similarity scores from vector queries to order retrieved documents. In parallel pipelines, each retrieval system independently scores candidates before combining outputs through fusion techniques such as Reciprocal Rank Fusion or Relative Score Fusion. This dual approach balances strict lexical constraints with conceptual relevance while mitigating individual retrieval blind spots.
Architectural mechanisms of hybrid retrieval
Lexical search retrieves documents by matching query terms against an inverted index that maps terms to the documents containing them [1][2]. In engines such as Apache Lucene and MongoDB Atlas Search, BM25 serves as the default scoring algorithm for keyword relevance [3]. BM25 scores documents based on term frequency saturation, which prevents keyword-stuffed documents from dominating by ensuring each additional occurrence has diminishing impact [4]. In addition, BM25 applies document length normalization relative to the average collection document length so a concise answer can outrank a lengthy text [5]. It also incorporates inverse document frequency, giving higher weight to rare terms than common terms [6]. Conversely, dense vector search represents text as numerical vectors to evaluate semantic similarity using metrics such as cosine similarity, Euclidean distance, or dot product [7][8].
Relying exclusively on one retrieval approach exposes blind spots, as full-text search misses relevant documents when wording changes, while vector search can be unreliable for exact identifiers [9]. Standard hybrid retrieval pipelines address these issues by executing a multi-stage process where metadata filters run first, full-text and vector retrievers run in parallel to generate two separate ranked lists with scores, and the outputs are merged and reranked [10][11]. Alternatively, full-text search can be used to narrow candidate sets by keyword, phrase, or metadata fields before running vector search inside that smaller set to rank by meaning [12].
Score fusion and ranking algorithms
When parallel retrieval executes, hybrid systems generate separate ranked lists with scores from keyword search and vector search [11]. In MongoDB Atlas, reciprocal rank fusion and relative score fusion are recommended techniques for combining BM25 and vector search scores into a unified ranking [8][13]. Relative score fusion normalizes scores from each search method, scaling them to a common range, typically 0 to 1, before combining them [14]. This normalization ensures that the relative importance of each search method is preserved even when original score distributions differ [14].
Reciprocal rank fusion calculates the reciprocal rank of each document across different search methods and combines these ranks into a unified score for each document [15]. An RRF algorithm calculates the reciprocal rank score as 1/(rank + k), where rank is the position of the document in the list and k is a constant that balances the influence of individual rankings and decides sensitivity to rank positions [16][17]. Experimental results demonstrate that RRF performs best when k is set to a small value, such as 60 [17]. The engine then sums the reciprocal rank scores obtained from each strategy to produce a single combined score, reranking documents so those with higher scores appear at the top of the final list [18]. In Azure AI Search, RRF merges and homogenizes rankings into a single result set whenever two or more queries execute in parallel, such as hybrid queries combining full-text and vector queries [19].
In Azure AI Search, full-text search scored with BM25 has no upper limit [20]. Vector search scored with HNSW yields score ranges of 0.333 to 1.00 for cosine similarity, and 0 to 1 for Euclidean and dot product metrics [7]. For hybrid search using RRF, the upper limit is bounded by the number of queries being fused, with each query contributing a maximum of approximately 1/k to the score [21]. Systems can assign different weights to search strategies or vector queries before combining scores to increase or decrease their importance [22][23].
Multilingual retrieval and secondary reranking
Hybrid retrieval workflows often combine scores and follow with reranking [11]. In Azure AI Search, semantic ranking can occur after RRF merging of results, returning a separate reranker score between 0.00 and 4.00 in the query response [24].
Research investigating embedding APIs on standard retrieval benchmarks, BEIR and MIRACL, demonstrates that ranking performance depends on target languages [25][26][27]. On English retrieval, reranking BM25 results using embedding APIs provides a budget-friendly approach that is most effective compared to employing embeddings as first-stage retrievers [25]. For non-English retrieval, reranking improves results, but a hybrid model combined with BM25 achieves the best performance, albeit at a higher cost [26].
Key facts
- BM25 calculates relevance using term frequency saturation, document length normalization, and inverse document frequency [4][5][6].
- Standard hybrid retrieval pipelines run metadata filters first, execute full-text and vector retrievers in parallel, and merge and rerank results [10][11].
- Full-text search misses relevant documents when wording changes, while vector search can be unreliable for exact identifiers [9].
- Reciprocal rank fusion calculates scores using 1/(rank + k), where rank is document position and k is a constant that performs best when set to a small value such as 60 [16][17].
- In RRF, reciprocal rank scores obtained from each strategy are summed to create a single score, placing higher-scoring documents at the top of the final ranking [18].
- Relative score fusion normalizes scores from each method to a common range, typically 0 to 1, preserving relative importance across different score distributions [14].
- In Azure AI Search, BM25 scores have no upper limit, vector similarity scores range from 0.333 to 1.00 for cosine and 0 to 1 for Euclidean and dot product, and RRF scores are bounded by the number of queries fused with each contributing at most approximately 1/k [20][7][21].
- Different retrieval strategies and vector queries can be weighted to adjust their influence before combining scores [22][23].
- Semantic ranking can take place after RRF merging in Azure AI Search, producing a separate score between 0.00 and 4.00 [24].
- Benchmark evaluations on BEIR and MIRACL show that reranking BM25 results is most effective for English, while a hybrid model with BM25 performs best for non-English retrieval [25][26][27].
Sources
-
Full-text search for RAG apps: BM25 & hybrid search redis.io
- [1]
Instead of scanning every document from top to bottom, it uses an inverted index, a data structure that maps terms to the documents containing them, so lookups are fast even across millions of records.
- [4]
BM25 scores documents based on three factors: Term frequency (TF) saturation: a word appearing 100 times doesn't score 100x higher than one appearance. Each additional occurrence has diminishing impact, which prevents keyword-stuffed documents from dominating results.
- [5]
Document length normalization: longer documents don't automatically win just because they contain more words. BM25 adjusts scores relative to the average document length in your collection, so a concise answer can outrank a lengthy spec.
- [6]
Inverse document frequency (IDF): rare terms carry more weight than common ones.
- [9]
Each approach has a blind spot. Full-text search misses relevant documents when wording changes ("terminate" vs. "cancel"). Vector search can be unreliable for exact identifiers when the token itself is the point
- [10]
The overall shape of most hybrid retrieval pipelines is similar: metadata filters first, then two retrievers (full-text and vector) running in parallel, then a merge and re-rank.
- [12]
Use full-text search to narrow the candidate set (by keyword, phrase, or metadata fields). Run vector search inside that smaller set to rank by meaning.
- [1]
-
Hybrid Search Explained With An In-Depth Guide www.mongodb.com
- [2]
Lexical search, sometimes referred to as keyword or full-text search, retrieves documents by matching exact terms or phrases in a query to the terms stored in an index.
- [3]
BM25 is the default scoring algorithm in Apache Lucene and MongoDB Atlas Search. It focuses on keyword relevance and ranks documents based on the frequency in which the queried keywords appear, considering factors like document length and overall term frequency.
- [8]
In MongoDB Atlas Vector Search, semantic similarity can be determined using Euclidean, cosine, or dot product metrics. Finally, the BM25 and vector search scores are combined to create a unified ranking, delivering the highest-ranked results to end users.
- [13]
There are multiple ways to combine the BM25 and vector search scores, among them are reciprocal rank fusion (RRF) and relative score fusion (RSF)—both of which are recommended techniques for hybrid search in MongoDB Atlas.
- [14]
RSF normalizes the scores from each search method, scaling them to a common range (typically 0 to 1), before combining them. This normalization ensures that the relative importance of each search method is preserved, even if their original score distributions differ.
- [15]
RRF works by calculating the reciprocal rank of each document across different search methods and then combining these ranks into a unified score for each document.
- [2]
-
Hybrid Search Scoring (RRF) - Azure AI Search learn.microsoft.com
- [7]
vector search @search.score HNSW algorithm, using the similarity metric specified in the HNSW configuration. 0.333 - 1.00 (Cosine), 0 to 1 for Euclidean and DotProduct.
- [17]
The score is calculated as 1/(rank + k), where rank is the position of the document in the list and k is a constant. Experiments show the algorithm performs best when you set k to a small value, such as 60.
- [19]
In Azure AI Search, RRF is used when two or more queries execute in parallel. Namely, for hybrid queries and for multiple vector queries. Each individual query produces a ranked result set, and RRF merges and homogenizes the rankings into a single result set for the query response.
- [20]
full-text search @search.score BM25 algorithm No upper limit.
- [21]
hybrid search @search.score RRF algorithm Upper limit is bounded by the number of queries being fused, with each query contributing a maximum of approximately 1/k to the RRF score
- [23]
You can also weight vector queries to increase or decrease their importance in a hybrid query.
- [24]
semantic ranking @search.rerankerScore Semantic ranking 0.00 - 4.00 Semantic ranking occurs after RRF merging of results. Its score (@search.rerankerScore) is always reported separately in the query response.
- [7]
-
Better RAG Results With Reciprocal Rank Fusion www.mongodb.com
- [11]
Since hybrid search is used, two separate ranked lists with scores are generated—one from keyword search and another from vector search. These scores are then combined, typically using a fusion technique such as reciprocal rank fusion, followed by reranking.
- [16]
Score = 1/(rank+k) k is a constant that helps in balancing the influence of individual rankings. The value of k decides the sensitivity to rank positions. In the above formula, rank is the position of the document in the list.
- [18]
The next step is to combine the scores obtained from each strategy and sum them to obtain a single score. Then, rank the documents again (rerank), based on the combined score. Documents having higher scores will be placed on the top in the final ranking.
- [22]
In some implementations, different search strategies can be assigned different weights before combining their scores. This allows the system to favor one retrieval method—such as semantic search—over another, depending on the use case or domain requirements.
- [11]
-
Evaluating Embedding APIs for Information Retrieval aclanthology.org
- [25]
We find that re-ranking BM25 results using the APIs is a budget-friendly approach and is most effective on English, in contrast to the standard practice, i.e., employing them as first-stage retrievers.
- [26]
For non-English retrieval, re-ranking still improves the results, but a hybrid model with BM25 works best albeit at a higher cost.
- [27]
Specifically, we wish to investigate the capabilities of existing APIs on domain generalization and multilingual retrieval. For this purpose, we evaluate the embedding APIs on two standard benchmarks, BEIR, and MIRACL.
- [25]