Performance of Unsupervised Learning Algorithms for Online Document Clustering
Dilip Singh Sisodia, Akanksha Verma · 2018 International Conference on Inventive Research in Computing Applications (ICIRCA) · 2018
The huge amounts of documents are available over the Internet. The effective automated process for indexing and searching of the online documents is essential for better user experience. The unsupervised learning techniques are useful for categorizing the available documents into correlated clusters for easier access. In this paper, partition based and hierarchical clustering techniques are discussed for clustering of the 20NewsGroups dataset documents. The experiments are performed using two partitions based on clustering such as K-Means and K-Medoids algorithms, and two hierarchical clustering, such as Single link analysis and complete link analysis methods. The different similarity measures such as Euclidean distance, Jaccard Coefficient, Cosine similarity, Pearson Correlation and ChiSquared distance is used for matching the similarity between documents. The clustering performance is evaluated using internal measures including DB Index, Silhouette Index, and C-Index; and external measures including F-Measure, Jaccard Index, and Rand Index. The results suggested that the partition based clustering algorithms perform better than the hierarchical clustering algorithms.