Web-Document Retrieval by Genetic Learning of Importance Factors for HTML Tags.
Sun Kim, Byoung‐Tak Zhang · 2000
Abstract. In contrast to conventional documents, a Web document con-sists of a number of tags which provide hints on the structure of the docu-ments. In this paper, we propose a Web-document retrieval method using the characteristics of HTML tags. This method learns the importance of tags from a training text set. We use a genetic algorithm for learning the importance weights. We also present a modied similarity measure which uses the tag information. Experiments have been performed on the TREC document collection consisting of 247,491 documents. Compared to the traditional IR method, the proposed method has achieved 15% improvement in average precision. 1