A Self-supervised Joint Training Framework for Document Reranking

Xiaozhi Zhu, Tianyong Hao, Sijie Cheng, Fu Lee Wang, Hai Liu · Findings of the Association for Computational Linguistics: NAACL 2022 · 2022

Pretrained language models such as BERT have been successfully applied to a wide range of natural language processing tasks and also achieved impressive performance in document reranking tasks.Recent works indicate that further pretraining the language models on the task-specific datasets before fine-tuning helps improve reranking performance.However, the pre-training tasks like masked language model and next sentence prediction were based on the context of documents instead of encour aging the model to understand the content of queries in document reranking task.In this paper, we propose a new self-supervised joint training framework (SJTF) with a selfsupervised method called Masked Query Pre diction (MQP) to establish semantic relations between given queries and positive documents.The framework randomly masks a token of query and encode the masked query paired with positive documents, and use a linear layer as a decoder to predict the masked token.In addition, the MQP is used to jointly opti mize the models with supervised ranking ob jective during fine-tuning stage without an ex tra further pre-training stage.Extensive exper iments on the MS MARCO passage ranking and TREC Robust datasets show that models trained with our framework obtain significant improvements compared to original models.

Read the paper · More papers on PaperTik