Improved Text Clustering with Neighbors

Sri Lalitha Y., Aliseri Govardhan · International Journal of Data Mining & Knowledge Management Process · 2015

With ever increasing number of documents on web and other repositories, the task of organizing and categorizing these documents to the diverse need of the user by manual means is a complicated job, hence a machine learning technique named clustering is very useful.Text documents are clustered by pair wise similarity of documents with similarity measures like Cosine, Jaccard or Pearson.Best clustering results are seen when overlapping of terms in documents is less, that is, when clusters are distinguishable.Hence for this problem, to find document similarity we apply link and neighbor introduced in ROCK.Link specifies number of shared neighbors of a pair of documents.Significantly similar documents are called as neighbors.This work applies links and neighbors to Bisecting K-means clustering in identifying seed documents in the dataset, as a heuristic measure in choosing a cluster to be partitioned and as a means to find the number of partitions possible in the dataset.Our experiments on real-time datasets showed a significant improvement in terms of accuracy with minimum time.

Read the paper · More papers on PaperTik