Using weight-controlled token matching to extract data from HTML files

Yan Xu, Tok Wang Ling · 2005

Most of the data stored in HTML files on the Web are semistructured. Extracting data and packing them into semistructured data models has received a lot of attention recently. We introduce a method that generates wrappers automatically for HTML files. The wrapper is generated from labeled training examples. We use weight-controlled token matching to locate the delimiters of the data of interest to the users. A list of tokens near the data is evaluated and each token is given a weight. We define a list of tokens to be the delimiter if the tokens are so important that the sum of the weights is larger than a threshold. A prototype is designed and a GUI is used to help build wrappers and extract data from the Web. Our method requires a small number of training examples and is flexible enough to deal with missing and misordered items. Compared to other approaches that may be too restrictive, our approach tolerates small modifications of HTML files.

Read the paper · More papers on PaperTik