Challenging the Invisible Web
Lieming Huang · Technischen Universität Darmstadt · 2008
The revolution of the World Wide Web (WWW or Web) has set off the globalization of information publishing and access. Organizations, enterprises, and individuals produce and update data on the Web everyday. With the explosive growth of information on the WWW, it becomes more and more difficult for users to accurately find and completely retrieve what they want. Although there are hundreds of thousands of general-purpose and special-purpose search engines and search tools, most users still find it hard to retrieve information precisely. Moreover, considering the great amount of valuable information hidden in the Invisible Web that is generally inaccessible to traditional "crawlers", providing users with an effective and efficient tool for Web searching is necessary and urgent. First, this dissertation proposes an adaptive data model for meta-search engines (ADMIRE) that can be used to formally and meticulously describe the user interfaces and query capabilities of heterogeneous search engines on the Internet. Compared with related work, this model focuses more on the constraints between the terms, term modifiers, attribute order, and the impact of logical operators. Second, this dissertation presents a constraint-based query translation algorithm. When translating a query from a meta-search engine to a remote source, the mediator considers the function and position restrictions of terms, term modifiers and logical operators among the controls in the user interfaces to the underlying sources sufficiently, thus allowing the meta-search engine to utilize the query capabilities of the specific sources as far as possible. In addition, a two-phase query subsuming mechanism is put forward to compensate for the functional discrepancies between sources, in order to make a more accurate query translation. Furthermore, this dissertation presents a mechanism for constructing adaptive, dynamically generated user interfaces for meta-search engines based on the above-mentioned model. The concept of control constraint rules has been proposed and applied to the user interface construction. Depending on the state of interaction between users and system, such meta-search engines adapt their interfaces to the concrete user interfaces of differing kinds of search engines (Boolean model with differing syntax, vector-space/probabilistic model, natural language support, etc.), so as to overcome the constraints of heterogeneous search engines and utilize the functionality of the individual search engines as much as possible. Finally, this dissertation also tackles some issues on wrapper generation and result merging for Web information sources. The experiments show that an information integration system with an adaptive, dynamically generated user interface, coordinating the constraints among the heterogeneous sources, will greatly improve the effectiveness of integrated information searching, and will utilize the query capabilities of sources as far as possible. The adaptive meta-search engine architecture proposed in this dissertation has been applied to the information integration of scientific publications-oriented search engines. It can also be applied to other generic domains or specific domains of information integration, such as integrating all kinds of WWW search engines (or search tools) and online repositories with quite different user interfaces and query models. With the help of source wrapping tools, they can also be used to integrate queryable information sources delivering semi-structured or non-structured data, such as product catalogues, weather reports, software directories, and so on.