Search for near-duplicate texts in the linguistic corpus VepKar

Fedor Bykov, Андрей Анатольевич Крижановский, Fedor Bykov, Andrew Anatoliyevich Krizhanovsky · Proceedings of the Karelian Research Centre of the Russian Academy of Sciences · 2023

Developers of linguistic corpora need to spot and eliminate text duplicates. An overview of approaches to searching for near-duplicate texts in various corpora is presented in this article. An algorithm and a program for searching for nearduplicate texts (based on the number of common bigrams) have been developed. Experiments were carried out with texts from the Veps and Karelian Open Corpus VepKar. The program found 100 pairs of the most similar texts and offered them to an expert, who confirmed 42 cases to be duplicates. Three metrics of text similarity were considered. The metric that was the closest to the expert’s output in its pairwise text alignments was identified using Kendall’s rank distance. The newly developed program will be a useful tool for editors of the VepKar text corpus.

Read the paper · More papers on PaperTik