Telugu Paraphrase Detection Using Siamese Network
M. Shiva Rohith, Mothukuri Jaswanth Venkat, Pasumarthy Venkata Akhil, Mandiga Sahasra Sai Tarun, Deepa Gupta · 2022 13th International Conference on Computing Communication and Networking Technologies (ICCCNT) · 2022
Paraphrasing methods will generate, identify or extract the phrases or sentences that convey the same meaning. In this work, given a pair of sentences, they are labeled as Paraphrased or Non-paraphrased. The ability to detect similar sentences is essential for several applications like text summarization, question answering, and plagiarism detection. Simple Telugu sentences are employed to create a Telugu paraphrase dataset; after paraphrasing, they are labeled as Paraphrased pair or non-paraphrased pair, respectively. iNLTK word embedding has been used with the Siamese network for calculating the similarity score between the pairs of sentences. And different deep learning architectures like the average scoring, Long Short-Term Memory (LSTM,) and Bidirectional Long Short-Term Memory (Bi-LSTM) are implemented for obtaining the sentence embedding. Their comparative study has been presented in this work.