Unsupervisedincremental acquisition of a thematic corpus from the web

Florence Duclaye, François Yvon, O. Collin · 2004

We present a nearly unsupervised learning methodology for automatically acquiring a thematic corpus from the Web. Relying on a bootstrapping mechanism, our system starts with one single linguistic expression of a given target semantic relationship. It then samples the Web so as to progressively accumulate a corpus of potential examples of the same relationship. Sampling steps alternate with filtering steps, making it possible to keep the corpus thematically focused. The corpus is finally analysed to search for potential paraphrases of the initial expression of the semantic relationship. These paraphrases will eventually be used to improve our question-answering system. We focus on die learning aspect of the system and reports experimental results regarding the effectiveness of our filtering strategy.

Read the paper · More papers on PaperTik