DIMENSIONALITY REDUCTION TECHNIQUES FOR SEARCH RESULTS CLUSTERING

Yoshi Gotoh · 2004

Search results clustering is an attempt to automatically organise a linear list of document references returned by a search engine into a set of meaningful thematic categories. Such a clustered view helps the users to identify documents of interest more quickly. One search results clustering method is the description-comes-first approach, whereby using a dimensionality reduction technique a number of meaningful group labels are identified, which then determine the content of the actual clusters. The aim of this project was to compare how three different dimensionality reduction techniques would perform as parts of the description-comes-first method in terms of quality of clustering and computational efficiency. The evaluation stage was based on the standard merge-then-cluster model, in which we used the Open Directory Project web catalogue as a source of human-clustered document references. During the course of the project we implemented a number of dimensionality reduction techniques in Java and integrated them with our description-comes-first search results clustering algorithm. We also created a simple benchmarking application, which we used to gather data for further comparisons and analysis. Finally, we have chosen one dimensionality reduction technique that performed best both in terms of clustering quality and computational efficiency.

Read the paper · More papers on PaperTik