Adaptive information extraction from online documents
Zhenmei Gu · 2006
As text-based resources available online continue to grow explosively, there is ever-increasing need for automatically extracting useful information from these textual data. Information Extraction (IE) is the problem of extracting specific information from textual documents for generating structured summaries. The research in IE has evolved from the IE systems that were manually built with hand-crafted rules, to more recent ones that obtain extraction knowledge automatically from annotated texts. The advantage of using machine learning is that it is easier to adapt an IE system to different extraction tasks. Such adaptability is a key requirement for practical IE systems. In view of these, in this thesis we focus our work on the IE systems that can learn extraction knowledge from data, hence adaptive IE systems. Our work is centered around the same goal to improve the robustness of current IE systems and to ease the difficulty of adapting an IE system to different extraction tasks. We begin our journey by first giving a full examination of the naive Bayes IE model, a purely adaptive IE model. We first present a formal naive Bayes modelling for IE problems, in which the formulation problem existing in previous naive Bayes IE work is corrected. We then investigate the smoothing techniques in the context of our naive Bayes IE systems, which is essentially a general issue associated with any probabilistic model. In particular, we design an appropriate smoothing strategy to be used with naive Bayes IE systems in order to obtain more stable probability estimation from training. Our experimental results show that a good smoothing method is critical to the robustness of naive Bayes IE systems. By completing several aspects of naive Bayes modelling for IE, our resulting systems demonstrate significant performance improvement over the existing naive Bayes IE systems. As witnessed in most existing probabilistic systems, a natural evolution from naive Bayes models is to more advanced Hidden Markov Models (HMMs). Our work on HMM IE solves the extraction redundancy issue existing in current HMM based IE systems that uses HMMs to model an entire document. Toward this end, we consider a segment-based HMM IE approach, in which a segment retrieval step is added for identifying extraction related segments from the whole document. (Abstract shortened by UMI.)