Document Clustering using 3-tuples.
Rajaraman Kanagasabai, Hong Pan · 2000
. We address the document clustering problem using a new representation called 3-tuples. We describe the 3-tuples representation and define a novel document similarity measure based on this representation. Using this measure, a new document clustering algorithm is proposed. We analyze the performance of our algorithm empirically using an entropy based cluster evaluation setup. We also present a comparative study with existing methods. 1 Introduction Document clustering is a method of grouping similar documents together from a given collection. It is assumed that no a priori knowledge on the grouping is known. This aspect differentiates it from Document Classification where the class/category labels are provided. Document clustering has applications in many scenarios such as: organizing the query-search results, navigating large text collections and event tracking. Various types of clustering algorithms have been discussed in the literature[9, 13, 4]. Our approach is based on a...