A Unified Framework for Tuning Hyperparameters in Clustering Problems
Xinjie Fan, Y. X. Rachel Wang, Purnamrita Sarkar, Yuguang Yue · Statistica Sinica · 2022
Selecting hyperparameters for unsupervised learning problems is challenging in general due to the lack of ground truth for validation.Despite the prevalence of this issue in statistics and machine learning, especially in clustering problems, there are not many methods for tuning these hyperparameters with theoretical guarantees.In this paper, we provide a framework relying on maximizing a trace criterion connecting a similarity matrix with clustering solutions, which has provable guarantees for selecting hyperparameters in a number of distinct models.We consider both the sub-gaussian mixture model and network models to serve as examples of i.i.d. and non-i.i.d.data.We demonstrate that the same framework can be used to choose the Lagrange multipliers of penalty terms in semidefinite programming (SDP) relaxations for community detection, and the bandwidth parameter for constructing kernel similarity matrices for spectral clustering.By incorporating a cross-validation procedure, we show the framework can also do consistent model selection for network models.Using a variety 1 Equal contribution 2