EXTRACTING ATTRIBUTES AND THEIR VALUES FROM WEB PAGES

Minoru Yoshida, Kentaro Torisawa, Jun’ichi Tsujii · Series in machine perception and artificial intelligence · 2003

We propose a method for extracting at-tributes and their values from Web pages. Our method makes use of word distri-butions estimated from plain Web pages. The key idea is to estimate word distri-bution by consulting ontologies built from HTML tables. In a series of experiments, we show that estimated word distribu-tions are useful for extracting attributes and their values in various kinds of HTML representations other than tables. 1

Read the paper · More papers on PaperTik