Optimizing Cluster Methods: Combining K-Means with Hierarchical Techniques for Better Results
Diyah Ruswanti, Ichwan Joko Prayitno, Firdhaus Hari Saputra Al Haris · 2024
The K-Means algorithm in the clustering process has limitations when dealing with clusters that have undefined sizes and shapes. One way to overcome this weakness is by using another algorithm to determine the number of clusters. Hierarchical clustering has the ability to recognize clusters of various sizes and shapes, and it also provides a visual representation in the form of a dendrogram. The collaboration between Hierarchical clustering and K-Means, when applied in clustering the research topics of final-year students, resulted in a total of 10 clusters. The object is Informatics Study Program at Sahid Surakarta University with 143 data that have been cleaned and ready to cluster. The aim of this research is to determine the contribution of the hierarchical method when combined with the K-means algorithm for clustering. The data used consists of abstracts from students' final projects, where the data has a varied form with different sizes and shapes of abstracts. Let me know if you need further adjustments! This study uses the CRISP-DM (Cross-Industry Standard Process Model for Data Mining) technique and uses a combination of hierarchical clustering and K-means clustering to group the data. The cluster center and number of clusters are determined using the hierarchical clustering technique. Next, K-means clustering optimizes the centroid location by repeatedly calculating the centroid of each cluster. This calculation continues until the centroid value is stable or the iteration limit is reached. After K-means reaches a stable centroid, the centroid value is considered accurate. Combining the two will provide better clustering results. Based on the clustering results, 10 clusters were formed. The fourth cluster is a non-specific cluster. The quality of the cluster results is calculated using intra-cluster similarity and inter-cluster similarity, with an average value of 0.210 for inter-cluster similarity and 0.716 for intra-cluster similarity. The cluster value is included in the “good” category. As a reference, the best inter-cluster similarity value is 0 and the best intra cluster similarity value is 1.