Web Information Extraction Based on Generalized User Navigation Path

Junpei Tsutsui, Kimihito Ito, Hiroki Arimura · 2007

This paper studies information extraction from the Web. Our method combines the advantages of an approach based on HTML page navigation and an approach based on HTML wrapper induction. The keys of our method are construction of extraction tree from navigation records and the use of machine learning techniques for automatically constructing wrappers to achieve flexible information extraction.

Read the paper · More papers on PaperTik