Learning to extract information from large websites using sequential models

V. G. Vinod Vydiswaran, Sunita Sarawagi · Conference on Management of Data · 2005

We propose a new method of information extraction from large websites by learning the sequence of links that lead to a specific goal page on the website. Sample applications include finding computer science publications starting from university root pages and fetching addresses of companies on a web database. We model the website as a graph on a set of important states chosen via domain knowledge and train a Conditional Random Field (CRF) over it. The conditional exponential models of CRFs enable us to exploit a variety of features including keywords and patterns extracted from and around hyperlinks and HTML pages and any sequential orderings amongst states. Our technique provides two times better harvest rates than techniques used in generic focused crawlers.

Read the paper · More papers on PaperTik