HaVAT: Automatic Cluster Structure Assessment in Unlabeled data
Pagadala Krishna Murthy, Punit Rathore · 2024
Clustering algorithms often rely on the input of the desired number of clusters, denoted as "k," to partition the data. However, a crucial question arises: do the data truly exhibit clusters, and if so, how many? This is known as clustering tendency assessment. Many variants of VAT have been proposed recently, and this family of algorithms called Visual Assessment of Clustering Tendency (VAT) has gained popularity among researchers in various domains. These algorithms aim to visually estimate the cluster structure and the number of clusters in the input dataset by reordering the pairwise dissimilarity matrix and creating a grayscale image called the reordered dissimilarity matrix image (RDI). Dark blocks in the RDI, representing pixels with low dissimilarity values, visually indicate potential clusters within the data. Although the VAT family of algorithms has proven valuable for estimating clusters in diverse datasets, manually interpreting the output, particularly with overlapping clusters or complex geometries, can be challenging. In this paper, we propose HaVAT, a novel method based on the Hough transform, to automate the interpretation of the RDI. Additionally, HaVAT can automatically determine the cluster hierarchy and obtain the optimal partition as part of the automated assessment based on a novel scoring mechanism. Our experiments on various datasets demonstrate HaVAT’s effectiveness and superiority over state-of-the-art methods in estimating cluster structure in terms of hierarchy and number of clusters.