An Algorithm for Clustering Search Results Based on Query Relevance Analysis

Zhonghua Yu · Journal of Chinese Computer Systems · 2011

With the popularity of the Internet and the rapid growth in quantity of web pages,the Search Engine has become a primary way to acquire information from the Internet.However,current leading search engines often respond to users with a long one-dimensional list consisting of snippets of returned web pages and being displayed by pages.In order to find the needed information,the users must be patient with browsing the list.To further improve the efficiency and quality of the information acquisition,and reduce the labor intensity of users,problems about reorganizing and re-mining search results are introduced and investigated in the literature,and among them search result clustering attracts most attentions,becoming a hot spot in the research field.In this paper,after analyzing the shortages of the existing relevant algorithms,a label-driven clustering algorithm based on query relevance analysis is proposed.The algorithm at first analyzes the query relevance of every phrase and chooses the phrases strongly associated with the query as candidate cluster labels,then determines the corresponding relationship between snippets and candidate clusters through the labels,and finally sifts and merges the resulted candidate clusters based on the evaluation of the candidate clusters and their labels.The experimental results under the same environment demonstrate that the proposed algorithm is superior to the related work,and just needs fewer support information resources.

Read the paper · More papers on PaperTik