Hybrid Clustering Techniques for Optimizing Online Datasets Using Data Mining Techniques

Bhupender Singh Rawat, Animesh Srivastava, Gulbir Singh, Gautam Kumar, Shiv Ashish Dhondiyal · 2023

Data elements of the same group are used in the clustering procedure. The most common clustering algorithm is based on initially chosen centroids randomly, which is the K-mean. The K-means clustering algorithm comprises numerical and category attributes that suffer from significant difficulty in initializing the cluster centre for large datasets. This paper suggests a hybrid K-Harmonic clustering algorithm that tries to overcome the problem of cluster center initialization of data with enormous datasets and numbers. By arranging the data, It is possible to improve the K-Harmonic means clustering method and was randomly used earlier for mixed datasets, making a new hybrid algorithm for K-Harmonic means. A definition is presented for the cluster's hub and a distance measure by using K-Harmonic means clustering function. Trials were conducted with pure large mixed datasets and categorical datasets of online shopping malls. Results imply that the effectiveness of the suggested grouping technique is unaffected by the initialization difficulty of the cluster. It has been shown through comparative studies of complexity and accuracy utilizing additional clustering methods and a variety of datasets that the suggested approach yields superior clustering results.

Read the paper · More papers on PaperTik