Interactive Learning of HTML Wrappers Using Attribute Classification

Michal Ceresna · 2005

Abstract. Reviewing the current HTML wrapping systems, it is possible to recognise two mainstream categories. The first category are systems based on various machine learning techniques with lower roll-out and maintenance costs, but reaching worse results and usually being specialised to particular domains. The second category are systems which allow to build more complicated wrapping solutions. But here a human wrapper designer is required to build and maintain the wrappers. Therefore, it increases costs for acquiring of the information. In this paper we apply machine learning techniques to automate and simplify building of the HTML wrappers by the wrapper designer. We present a learning algorithm that creates wrappers from interaction with an human designer. She only marks positive and negative example instances inside of the currently rendered Web and an HTML wrapper is generated from this interaction. The learning algorithm presented in this paper is based on clustering of the example instances with respect to the similarity of their tree shape and on classification of HTML attributes inside of each cluster. 1

Read the paper · More papers on PaperTik