A Comparison-Based Soft Clustering Algorithm for Documents

International journal of research studies in computer science and engineering · 2019

A challenge in document clustering is that many documents contain multiple subjects.For instance, a Web page discussing the University of Memphis's research on wild tiger fits under the categories of both "wild animals" and "Memphis".Thus, the clustering algorithm should discover this and put the document under both clusters.This suggests the use of "soft clustering" algorithman algorithm that allows a document to appear in multiple clusters.This can help users to discover multiple themes in a documentby looking at the multiple clusters that a document belongs to.Soft clustering can also help form clusters containing combination of existing topics.For instance, we might want to have a separate cluster of documents about the University's research on wild tigers.This is possible if documents can fall into multiple clusters.Many soft clustering algorithms have been developed and most of them are based on the Expectation-Maximization (EM) algorithm [6].They assume an underlying probability model with parameters that describe the probability that an object belongs to a certain cluster.Based on the data given, the algorithms try to find the best estimation of the parameters.However, a drawback of such algorithms is that they tend to be computationally expensive.In this paper, we take a different approach.Instead of assuming an underlying probability model, we only assume that we are given a similarity function f (x, y), which given documents x and y, returns a Abstract: Data document clustering is an most important tool for searching document such as Web search engines.Clustering data documents enables the accessor to have a good overall view of the information contained in the documents that he has.However, existing clustering algorithms faces from various aspects; complex clustering algorithms (where each document belongs to exactly one cluster) cannot detect the multiple themes of a document, while flexible such as soft clustering algorithms (where each document can belong to multiple clusters) are usually inefficient.We propose CSCA (Comparison-based Soft Clustering), an efficient soft clustering algorithm based on a given similarity measure.CSCA requires only a similarity measure for clustering and uses randomization to help make the clustering efficient.Comparison with existing complex hard clustering algorithms like K-means and its variants shows that CSCA is both effective and efficient.

Read the paper · More papers on PaperTik