Detecting and Clustering Similar Results of Search Engine by Exploiting Web Page's Contents

Kai Gao, Huicong Wu · 2008

This paper presents an approach to detect and cluster similar results of search engine based on analyzing pages' URLs and their contents. A novel hash function, together with a Chinese key concept extractor module, has been used. The similar measurement on key concept overlap degree is proposed to cluster similar retrieval results. This can minimize the overlap effectively. The experimental results show the feasibility of the approach. On the basis of the above works, a search engine has been developed.

Read the paper · More papers on PaperTik