Clustering Algorithm with a Novel Similarity Measure

Gaddam Saidi Reddy · IOSR Journal of Computer Engineering · 2012

Clustering is one of the data mining and text mining techniques used to analyze datasets by dividing it into meaningful groups.The objects in the dataset can have certain relationships among them.All clustering algorithms assume this before they are applied to datasets.The existing algorithms for text mining make use of a single viewpoint for measuring similarity between objects.Their drawback is that the clusters can't exhibit the complete set of relationships among objects.To overcome this drawback, we propose a new similarity measure known as multi-viewpoint based similarity measure to ensure the clusters show all relationships among objects.We also proposed two clustering methods.The empirical study revealed that the hypothesis "multi-viewpoint similarity can bring about more informative relationships among objects and thus more meaningful clusters are formed" is proved to be correct and it can be used in the real time applications where text documents are to be searched or processed frequently.

Read the paper · More papers on PaperTik