Computational aspects of resilient data extraction from semistructured sources (extended abstract)
Hasan Davulcu, Guizhen Yang, Michael Kifer, I. V. Ramakrishnan · 2000
Automatic data extraction from semistructured sources such as HTML pages is rapidly growing into a problem of significant importance, spurred by the growing popularity of the so called “shopbots” that enable end users to compare prices of goods and other services at various web sites without having to manually browse and fill out forms at each one of these sites.