Cross Language Similarity Detection Technology Combining Optimized K-Means Algorithm with CL-ESA

Yuxin Gan, Qian Xu · 2024

Academic plagiarism is becoming an increasingly serious cross language problem, and the academic community is paying more and more attention to cross language text similarity detection algorithms. In order to improve the accuracy of cross linguistic text similarity detection algorithms, the study first combines word level precise translation and sentence level contextual understanding, and combines machine translation algorithms based on word granularity text with those based on sentence granularity text to design an improved machine translation algorithm. At the same time, a heuristic retrieval mechanism is introduced on the basis of CL-ESA to narrow down the search scope of cross linguistic text similarity detection, and an improved index document selection algorithm (IIDSA) is proposed in combination with the K-Means algorithm. Finally, a cross lingual text similarity detection model is constructed by combining the IIDSA algorithm with an improved machine translation algorithm. This study extracted 500 pairs of concept aligned Chinese and English texts from Wikipedia as experimental texts to test the algorithm. The results show that the accuracy of the model is 96.2%, which is superior to other models. The above conclusions verify that the proposed cross language similarity detection model can meet people's needs for cross language text similarity detection and to some extent promote the development of academic research.

Read the paper · More papers on PaperTik