A novel highly scalable clustering algorithm based on hyper edges and successive merging with randomization for complex data sets
Arindam Ghosh, Ruma Ghosh, Debaprasad Mukherjee · 2014
Classification through clustering of complex data sets is a fundamental problem in computer science, and it has various applications in the fields of biomedicai sciences, weather prediction, web search, security and surveillance, information retrieval etc. Several algorithms are available on this area. Here, we develop a new algorithm for clustering of graph structured vector valued data, for final classification of the data into proper meaningful groups. In this algorithm, hyperedges are created where sets of consecutive edges based on similarity of feature vectors of nodes are created and successively merged. Furthermore, several randomization steps have been incorporated to overcome any errors in the classification due to bias in the input data set. We have justified the validity, accuracy and performance of the algorithm through analysis and preliminary simulations. Our analysis and simulations indicate a satisfactory outcome and provide support to our inferences.