SIMILARITY BASED AUTOMATIC ITEM CLUSTERING FOR EFFICIENT CLASSIFICATION OF INFORMATION SPACE
M.N.Srinivas K.Shiva Reddy, Ravulapalli Lakshmi Tulasi, Iaeme Publication · 2013
Clustering is a useful technique that organizes a large quantity of unordered objects into a small number of meaningful and coherent clusters. The goal of the clustering is to assist in the location of information. It is an essential data analysis method used in many applications such as psychology, biology, information retrieval and mining technologies. Nowadays all manual documents are in automated form, because of fast access and lesser storage. So, to retrieve appropriate documents from huge database it is a major issue. Clustering documents to related groups is one of the active field of research in different fields of text mining, topic tracking systems, and question answering systems. We are proposing four eminent clustering algorithms that use standard similarity metrics on a document corpus to perform the clustering. Here, presents a survey on these existing document clustering algorithms and proposes a framework for comparing them using a similarity measure with respect to a number of documents and processing time.