Twitter Paraphrase Identification with Simple Overlap Features and SVMs

Asli Eyecioglu, Bill Keller · 2015

We present an approach to identifying Twitter paraphrases using simple lexical over-lap features. The work is part of ongoing re-search into the applicability of knowledge-lean techniques to paraphrase identification. We utilize features based on overlap of word and character n-grams and train support vector machine (SVM). Our results demonstrate that character and word level overlap features in combination can give performance comparable to methods employing more sophisticated NLP processing tools and external resources. We achieve the highest F-score for identifying paraphrases on the Twitter Paraphrase Corpus as part of the SemEval-2015 Task1.

Read the paper · More papers on PaperTik