A Paraphrase Identification Approach in Paragraph Length Texts
Arwa Al Saqaabi, Eleni C. Akrida, Alexandra Ioana Cristea, Craig D. Stewart · 2022 IEEE International Conference on Data Mining Workshops (ICDMW) · 2022
Measuring the semantic similarity of natural language is a fundamental issue in many tasks, such as paraphrase identification (PI) and plagiarism detection (PD) which are intended to solve maj or issues in education. Various approaches that have been suggested to tackle this issue, including machine learning (ML) and deep learning (DL) methods. Unlike prior research, where the focus has been on detecting paraphrases in short and sentence-level texts, we focus on the not yet explored area of paraphrase detection in paragraphs. We consider that the meaning of a piece of text can be broken into more than one sentence, which is over and beyond sentences as extracted from two benchmark datasets (Webis-CPC-ll and MSRP). We use TF-IDF, Bleu metric, N-gram overlap, and Word2vec as features, and then invoke SVM as a classifier. The contribution of this paper clearly indicates that, on a commonly used evaluation set, considering text at the length of a paragraph is more appropriate than considering short or long text for ML and DL approaches. Additionally, our method outperforms the existing work done on the Webis-CPC-ll dataset.