A Survey of Cross-Lingual Plagiarism Detection using Natural Language Processing

K. S. Aravind · International Journal for Research in Applied Science and Engineering Technology · 2020

Plagiarism detection is gaining importance due to requirements for integrity in Research works especially when it comes to Cross-lingual plagiarism. In this paper, we have researched a new approach for Cross-Lingual sentence level plagiarism detection. The proposed approach works as initially the suspicious document is translated to a particular language and candidate document is retrieved for further processing. Pre-processing techniques will be selected using NB Classifier for better efficiency which will be used on document which includes POS tagging, Chunking, Stop-word removal, Lemmatization, Syntactic-semantic rule. Then Machine Learning classification algorithms such as Support Vector Machine (SVM), Naïve Bayes Classifier (NBC), Decision Tree (DT) are used to check plagiarism. The Evaluation and Analysis are performed using F-score, Accuracy, Precision and Recall. Our aim in this paper is to find the plagiarism and improve the accuracy of plagiarism detection between two languages by using concepts of natural language processing.

Read the paper · More papers on PaperTik