Learning Query Token Importance for Effective Document Retrieval with Verbose Queries

Dipannita Podder, Jiaul H. Paik, Pabitra Mitra · 2024

Retrieving relevant documents with verbose queries is challenging due to the presence of extraneous terms.To address this, the centrality score of query terms is often estimated and integrated into traditional retrieval models.Recently, dense retrieval models have shown strong performance, where the relevance score is computed by analyzing the context of the document and query.Among these, multi-vector approaches such as Contextualized Late Interaction over BERT (ColBERT) embed each query token separately and treat all tokens with equal importance while estimating the relevance scores.Thus, the retrieval effectiveness of ColBERT degrades when the query contains extraneous terms.In this work, we propose a model that learns the importance of individual query tokens in verbose queries by leveraging the representations from the pretrained query encoder of ColBERT.During inference, the model assigns an importance score to each query token, which is then incorporated into the ranking function of ColBERT so that the token matching with more important tokens can be prioritized.We assign gold-standard labels for the tokens of training queries using an automatic annotation method leveraging the publicly available topics from NIST datasets.Experimental results demonstrate that the proposed method outperforms existing baselines across various test collections with verbose queries, and performs at par with the fine-tuned ColBERT, which is specifically fine-tuned for longer queries on each dataset.

Read the paper · More papers on PaperTik