Extracting Partial Structures from HTML Documents

Hiroshi Sakamoto, Yoshitsugu MURAKAMI, Hiroki Arimura, Setsuo Arikawa · Kyushu University Institutional Repository (QIR) (Kyushu University) · 2001

The new wrapper model for extractiong text data from HTML documents is introduced. The Kushmerick's wrapper class (Kusshmerick 2000) may be unsuccessful in the case that sufficiently long delimiters are not found. The wrapper class introduced in this paper partially overcomes this difficulty by using the tree structures of HTML documents. The learning problem to learn such a wrapper program from given text is considered. Moreover, we try to expand our wrapper to extract a portion of HTML not only text attributes.

Read the paper · More papers on PaperTik