Conceptual model of automatic system of near duplicates detection in electronic documents

Andrii O. Biloshchytskyi, Alexander Kuchansky, Svitlana Biloshchytska, Anastasiia Dubnytska · 2017

In the article, the conceptual model of near duplicates detection in electronic documents is considered. The model provides separation from the document of different data types (the text, the numerical sequences, images, diagrams and mathematical formulas) and applications to their analysis of the special tools allowing to identify similarities between fragments of the incoming document and documents which are stored in the database. This conceptual model can be used for establishment of loans in scientific publications and dissertation works for the purpose of plagiarism counteraction, and to detect duplicates in the descriptions of CAD projects.

Read the paper · More papers on PaperTik