Naive Bayes Modeling with Proper Smoothing for Information Extraction

Zhenmei Gu, N. Cercone · 2006

Information extraction (IE) summarizes a collection of documents into a structural representation by identifying specific facts from text. The naive Bayes model is one of the first statistical models that have been applied to IE for learning extraction patterns from labeled data. In spite of the simplicity and popularity of the naive Bayes model, we have observed a formulation problem in previous work on naive Bayes IE. In this paper, we present a formal naive Bayes modeling for IE, by which the derived formula for the filler probability estimation is more theoretically sound. We also address smoothing techniques in order to overcome the data sparseness problem. Our proposed smoothing strategy is shown to be critical to the robustness of a naive Bayes IE system. Experimental results show that our naive Bayes IE systems achieve better extraction performance compared to related work.

Read the paper · More papers on PaperTik