Siphoning Hidden-Web Data through Keyword-Based Interfaces: Retrospective
Luciano Barbosa, Juliana Freire · Cadernos de Linguística e Teoria da Literatura (Universidade Federal de Minas Gerais) · 2010
In this paper, we proposed the first, fully-automatic approach to crawling the Hidden Web throughkeyword-based interfaces. Our crawler uses an algorithm for automatically deriving a series ofkeyword-based queries whose goal is to obtain high coverage while minimizing the costs. In otherwords, our goal is to retrieve as much of the hidden contents as possible while minimizing the numberof required queries. The intuition behind our algorithm is that, by obtaining samples of the hiddencontents in a online database or document collection, we are able to discover keywords that have highfrequency. Then, by using these high-frequency keywords we are able to construct queries that returna large number of answers.