Explaining BERT model decisions for near-duplicate news article detection based on named entity recognition
A. Stockem, Fatih Gedikli · 2023
While it is a straight-forward task to detect lexically identical documents, it is challenging to do this on documents that are semantically similar. This work deals with the detection of near-duplicate news articles on various websites. The detection of documents of related content is especially interesting for search engines and recommender systems. Search engines aim at eliminating duplicates to prevent the user from seeing the same document in different versions. Recommender systems filter out items that are too similar to those the user already knows to tackle the problem of overspecialization. A near-duplicate is in some cases a mere paraphrasing of a previous article, in other cases it contains corrections or added content as a follow-up article. In this work, we investigate, among other things, the extent to which named entities such as people, places, and organisations that appear in the article text can be used to detect near-duplicates. Evaluating a fine-tuned BERT model on common named entities in an article pair, performance measures greater than 97 % are achieved. The SHAP library is then used to get insight into the model decisions and feature importance, which is the individual words in a text document. Overall, the results show that named entities are very suitable for the detection of near-duplicates. Furthermore, the explanations are in agreement with an intuitive human interpretation.