Adding Explainability to Visual Clustering Tendency Assessment (VAT) Methods Through Feature Importance

B Srinath Achary, Rankit Kachroo, Punit Rathore · 2024

Clustering is an important unsupervised learning approach to discover interesting patterns from the dataset. A preliminary problem in clustering involves determining whether and how many clusters exist in input data – also known as the clustering tendency assessment problem. Recently, a family of algorithms called Visual Assessment of Clustering Tendency (VAT) has gained popularity among researchers from various domains to visually estimate the number of clusters from a reordered dissimilarity matrix image (RDI) of the input data. However, just knowing the number of clusters and the assignment of data points to these clusters is insufficient from an end-user perspective. In this paper, to enhance the interpretability and explainability of VAT methods, we introduce a novel methodology, called FIM-VAT, to identify the importance of individual features in a dataset concerning the corresponding VAT output (RDI) by leveraging the concept of Spearman rank correlation. Experimental validation on eleven datasets, both real and synthetic, demonstrates the performance of our proposed technique in identifying feature importances in VAT for cluster structure assessment.

Read the paper · More papers on PaperTik