EXTRACTING ATTRIBUTES AND THEIR VALUES FROM WEB PAGES
Minoru Yoshida, Kentaro Torisawa, Jun’ichi Tsujii · Series in machine perception and artificial intelligence · 2003
We propose a method for extracting at-tributes and their values from Web pages. Our method makes use of word distri-butions estimated from plain Web pages. The key idea is to estimate word distri-bution by consulting ontologies built from HTML tables. In a series of experiments, we show that estimated word distribu-tions are useful for extracting attributes and their values in various kinds of HTML representations other than tables. 1