An improved method in clustering Web retrieval result based on relevance feedback

Xinye Li · 2011

Since the number of Web retrieval result is very large, the performance and reasonableness of clustering Web retrieval result are important. Existed methods cost much time while clustering all retrieval result and there were many unrelated document in their clustering result. To avoid the disadvantage, this paper proposed an improved k-means algorithm by using a few of related and unrelated feedback to guide clustering Web retrieval result. The improved algorithm first selected initial cluster metroid based on feedback messages, then during the clustering process, it removed large unrelated documents which increased the clustering speed and optimized the clustering result. During the clustering process, the metroids of clusters including unrelated documents needn't be modified in order to avoid noise influence. Experiment result illustrate that our algorithm is superior to the traditional k-means algorithm.

Read the paper · More papers on PaperTik