Machine learning based paraphrase identification system using lexical syntactic features

Rutal S. Mahajan, Mukesh A. Zaveri · 2016

During the natural language communication, meaning understanding is the complex task that humans learn from their childhood but to automate this process of meaning understanding for computers has great real world applications. Simple text processing tasks are not enough to uncover the meaning from given unstructured natural language text. Our current research focuses on the issues pertaining to the same. Paraphrase identification is such important task of identifying the meaning similarity between two text segments in natural language understanding system. Proposed a machine learning system uses lexical features and dependency based features for sentence level paraphrase identification. The performance of proposed system is evaluated by conducting experiment on standard Microsoft paraphrase corpus. Moreover, a comparative study of current system with other machine learning based systems on Microsoft paraphrase corpus for paraphrase identification is carried out. The proposed system achieves competitive results compare to other state-of-the art machine learning systems by using simple linguistic features. The system using SVM classifier achieves 81.41% f-score by using simple lexical features only. Voting based classifier scores 80.97% with lexical features. Results with dependency features are highly sensitive to minor syntactic change.

Read the paper · More papers on PaperTik