First European Web Mining Forum

Bettina Berendt, Andreas Hotho, Maarten van Someren, Myra Spiliopoulou, Gerd Stumme · 2003

Abstract. This paper presents a novel method for extracting information from collections of Web pages across different sites. Our method uses a standard wrapper induction algorithm and exploits named entity information. We introduce the idea of post-processing the extraction results for resolving ambiguous facts and improve the overall extraction performance. Post-processing involves the exploitation of two additional sources of information: fact transition probabilities, based on a trained bigram model, and confidence probabilities, estimated for each fact by the wrapper induction system. A multiplicative model that is based on the product of those two probabilities is also considered for post-processing. Experiments were conducted on pages describing laptop products, collected from many different sites and in four different languages. The results highlight the effectiveness of our approach. 1 Introduction Wrapper induction (WI) [7] aims to generate extraction rules, called wrappers

Read the paper · More papers on PaperTik