On Mining DOM Trees to build Information Extractors

Gretel Fernández, Hassan A. Sleiman, Rafael Corchuelo, Rafael Zancan Frantz · Deposito de Investigacion Universidad de Sevilla (University of Seville) · 2011

Abstract—The Web is the largest information repository. The information it contains is usually available in human-friendly formats. Companies are interested in using this information. The problem is that they need it in structured formats so that they can use it in automated business processes. In the literature, there are many proposals to infer information extractors. They build on machine learning techniques that attempt to infer a pattern in the HTML or XPath sources. To the best of our knowledge, no-one has ever explored using datamining techniques on DOM trees. In this paper, we report on a methodology that builds on datamining CSS features and a few other DOM features. Our results prove that this methodology is promising. Index Terms—Datamining Techniques, Information Extractor, Machine Learning

Read the paper · More papers on PaperTik