An Evolutionary Model for Measuring Document Relevance in a Focused Web Spider
I.F.V. Lopez, Pavel Anselmo Álvarez Carrillo, Eduardo Fernández-González · 2008
Exploring the Web in search of relevant information is a difficult task due to the vast amount of documents it stores and to the heterogeneity of such documents. Using automated systems such as search engines help users cope with the size of the Web. However the results produced by these systems usually contain documents from a large variety of topics with little or no relevance to the end user. In this work, we propose a model that can be used by a Web spider to selectively explore the Web for relevant documents. In this model, two criteria are used for assessing document relevance; content and structure. These two criteria are integrated in a fuzzy predicate that indicates the degree of relevance of a document with respect to a user-defined topic. The parameters of the proposed model are generated by a genetic algorithm that solves a bi-criteria optimization problem.