A Task-specific Approach for Crawling the Deep Web.
Manuel Álvarez, Juan Raposo, Fidel Cacheda, Alberto Pan-Bermudez · 2006
Abstract — There is a great amount of valuable information on the web that cannot be accessed by conventional crawler engines. This portion of the web is usually known as the Deep Web or the Hidden Web. Most probably, the information of highest value contained in the deep web, is that behind web forms. In this paper, we describe a prototype hidden-web crawler able to access such content. Our approach is based on providing the crawler with a set of domain definitions, each one describing a specific data-collecting task. The crawler uses these descriptions to identify relevant query forms and to learn to execute queries on them. We have tested our techniques for several real world tasks, obtaining a high degree of effectiveness. Index Terms—Crawler, Deep Web, HTML Forms, Server-Side. I.