Matching Long-form Document with Topic Extraction and Aggregation

Hua Cai, Jingxi Hu, Ren Ma, Yixiao Lu, Qing Xu · 2022

BERT-based models have been widely used for document matching, but they generally do not perform well on the matching of long-form documents, as the sequence length limitation could lead to loss of information in the document. Also, the increased noise of a long-form document would further complicate the capture of key matching signals. To tackle these existing problems, we propose a new long-form document matching model, named EA-BERT. In this model, a set of topics are first extracted from the pair of documents. For each topic, we gather related sentences from both documents to form a "bag of sentences", which is then encoded by BERT into a vector containing topic-level information. All topic vectors are then passed to a transformer encoder to aggregate topic-level information into a document-level matching result for the pair of documents. In this way, our approach overcomes the sequence length limitation of BERT-based models, filters out the increased noise of long-form documents, and integrates matching signals from multiple topics. Experiments show that the proposed method outperforms current long-form document matching models on both Chinese and English document-level datasets.

Read the paper · More papers on PaperTik