Labeling Clusters of Search Results
Martin Nycander · KTH Publication Database DiVA (KTH Royal Institute of Technology) · 2013
This project evaluates different algorithms which could be used in a summary information retrieval (IR) application for Swedish texts. Instead of the traditional search results the summary application would generate a summary document of the various subtopics of an IR query. First it is noted that in order to nd subtopics of a query, some kind of document clustering is needed. k-means is chosen as a candidate document clustering algorithm and evaluated in the environment of an IR application. It is found to be fast enough and to work better than the random clustering algorithm. Although it is argued that it is not good enough to be used in a summary/labeling context. Secondly the project looks into labeling algorithms to be used in the aforementioned IR application. Four algorithms were evaluated: TF centroid, TF- IDF centroid, mutual information and CorePhrase. None were deemed to generate high enough quality labels to be useful, but it was noted that TF-IDF is a good choice when constructing labels from an IR query.