Detection of Fuzzy Text Duplicates in Large Amounts of Information
Ekaterina V. Sharapova, Ruslan V. Sharapov · 2019
Fuzzy duplicates are texts that have undergone some adjustments and modifications, but retain a semantic similarity with the originals. The problem of detecting texts that are significantly similar to each other (fuzzy duplicates) arises when building plagiarism search systems, news aggregators and other intelligent systems. Currently, a number of methods have been developed to detect such documents. However, these methods allow either to quickly detect similar documents with a decrease in accuracy, or to search with high accuracy but for an unacceptably long time. The work is devoted to the development of methods for fuzzy duplicates detection that provide work with large amounts of information with minimal time and high search accuracy.