A Hybrid Approach to Paraphrase Detection Based on Text Similarities and Machine Learning Classifiers
Mena Hany, Wael Hassan Gomaa · 2022 2nd International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC) · 2022
In the realm of natural language processing (NLP), paraphrase detection is a highly common and significant activity. Because it is involved in a lot of complicated and complex NLP applications like information retrieval, text mining, and plagiarism detection. The proposed model finds the best combination of the three types of similarity techniques that are string similarity, semantic similarity and embedding similarity. Then, inputs these similarity scores that range from 0 to 1, to the machine learning classifiers. This proposed model will be benchmarked on “the Microsoft research paraphrase corpus” dataset (MSRP) and from this approach for paraphrase detection problem, the accuracy acquired is 75.78% and F1-Score of 83.01 %.