Expanding textual entailment corpora fromWikipedia using co-training

Fabio Massimo Zanzotto, Marco Pennacchiotti · International Conference on Computational Linguistics · 2010

In this paper we propose a novel method to automatically extract large textual entailment datasets homogeneous to existing ones. The key idea is the combination of two intuitions: (1) the use of Wikipedia to extract a large set of textual entailment pairs; (2) the application of semisupervised machine learning methods to make the extracted dataset homogeneous to the existing ones. We report empirical evidence that our method successfully expands existing textual entailment corpora.

Read the paper · More papers on PaperTik