Learning to Recognize Ancillary Information for Automatic Paraphrase Identification

Simone Filice, Alessandro Moschitti · 2016

Previous work on Automatic Paraphrase Identification (PI) is mainly based on modeling text similarity between two sentences.In contrast, we study methods for automatically detecting whether a text fragment only appearing in a sentence of the evaluated sentence pair is important or ancillary information with respect to the paraphrase identification task.Engineering features for this new task is rather difficult, thus, we approach the problem by representing text with syntactic structures and applying tree kernels on them.The results show that the accuracy of our automatic Ancillary Text Classifier (ATC) is promising, i.e., 68.6%, and its output can be used to improve the state of the art in PI.

Read the paper · More papers on PaperTik