DeBEIR: A Python Package for Dense Bi-Encoder Information Retrieval

Vincent Nguyen, Sarvnaz Karimi, Zhenchang Xing · The Journal of Open Source Software · 2023

Information Retrieval (IR) is the task of retrieving documents given a query or information need.These documents are retrieved and ranked based on a relevance function or relevance model such as Best-Matching 25 (BM25) (Robertson et al., 1995).Although deep learning has been successful in other computer science fields, such as computer vision with AlexNet (Krizhevsky et al., 2012) and Inception (Szegedy et al., 2014) and natural language processing with transformers (Devlin et al., 2019;Lee et al., 2019;Yang Liu & Lapata, 2019); success in information retrieval was limited due to comparisons against weak baselines (Yang et al., 2019).However, in 2019 (Lin, 2019), deep learning in information retrieval could surpass less computationally intensive keyword-based statistical models in terms of retrieval effectiveness, sparking a resurgence in the field of dense retrieval.Dense retrieval is the task of retrieving documents given a query or information need using a dense vector representation of the query and documents (Lin et al., 2021).The dense vector representation is obtained by passing the query and documents through a neural network.The neural network is usually a pre-trained language model such as BERT (Devlin et al., 2019) or RoBERTa (Yinhan Liu et al., 2019).The dense query vector representation is then used to retrieve documents using a similarity function such as cosine similarity.

Read the paper · More papers on PaperTik