Chunking based malayalam paraphrase identification using unfolding recursive autoencoders

R. Praveena, M Anand Kumar, K. P. Soman · 2017

Paraphrase Detection is the task of examining if two sentences convey the same meaning or not. Here, in this paper, we have chosen a sentence embedding by unsupervised RAE vectors for capturing syntactic as well as semantic information. The RAEs learn features from the nodes of the parse tree and chunk information along with unsupervised word embedding. These learnt features are used for measuring phrase wise similarity between two sentences. Since sentences are of varying length, we use dynamic pooling for getting a fixed sized representation for sentences. This fixed sized sentence representation is the input to the classifier. The DPIL (Detecting Paraphrases in Indian Languages) dataset is used for paraphrase identification here. Initially, paraphrase identification is defined as a 2-class problem and then later, it is extended to a 3-class problem. Word2vec and Glove embedding techniques producing 100, 200 and 300 dimensional vectors are used to check variation in accuracies. The baseline system accuracy obtained using word2vec for 2-class problem is 77.67% and the same for 3-class problem is 66.07%. Glove gave an accuracy of 77.33% for 2-class and 65.42% for 3-classproblem. The results are also compared with the existing open source word embedding and our system using Word2vec embedding is found to outperform better. This is a first attempt using chunking based approach for identification of Malayalam paraphrases.

Read the paper · More papers on PaperTik