ROB: Using Semantic Meaning to Recognize Paraphrases
Rob van der Goot, Gertjan van Noord · 2015
Paraphrase recognition is the task of identifying whether two pieces of natural language represent similar meanings.This paper describes a system participating in the shared task 1 of SemEval 2015, which is about paraphrase detection and semantic similarity in twitter.Our approach is to exploit semantically meaningful features to detect paraphrases.An existing state-of-the-art model for predicting semantic similarity is adapted to this task.A wide variety of features is used, ranging from different types of models, to lexical overlap and synset overlap.A maximum entropy classifier is then trained on these features.In addition to the detection of paraphrases, a similarity score is also predicted, using the probabilities of the classifier.To improve the results, normalization is used as preprocessing step.Our final system achieves a F1 score of 0.620 (10th out of 18 teams), and a Pearson correlation of 0.515 (6th out of 13 teams).