Expanding textual entailment corpora fromWikipedia using co-training
Fabio Massimo Zanzotto, Marco Pennacchiotti · International Conference on Computational Linguistics · 2010
In this paper we propose a novel method to automatically extract large textual entailment datasets homogeneous to existing ones. The key idea is the combination of two intuitions: (1) the use of Wikipedia to extract a large set of textual entailment pairs; (2) the application of semisupervised machine learning methods to make the extracted dataset homogeneous to the existing ones. We report empirical evidence that our method successfully expands existing textual entailment corpora.