Conceptual model of automatic system of near duplicates detection in electronic documents
Andrii O. Biloshchytskyi, Alexander Kuchansky, Svitlana Biloshchytska, Anastasiia Dubnytska · 2017
In the article, the conceptual model of near duplicates detection in electronic documents is considered. The model provides separation from the document of different data types (the text, the numerical sequences, images, diagrams and mathematical formulas) and applications to their analysis of the special tools allowing to identify similarities between fragments of the incoming document and documents which are stored in the database. This conceptual model can be used for establishment of loans in scientific publications and dissertation works for the purpose of plagiarism counteraction, and to detect duplicates in the descriptions of CAD projects.