k-POD: A Method for Clustering Partially Observed Data
Jocelyn T. Chi, Eric C. Chi, Richard G. Baraniuk · arXiv (Cornell University) · 2014
The clustering problem is ubiquitous in exploratory data analysis, yet surprisingly little attention has been given to the problem when data are missing. Mainstream approaches to clustering partially observed data focus on reducing the missing data problem to a completely observed formulation by either deleting or imputing missing data. These incur costs, however. Deletion discards information, while imputation requires strong assumptions about the missingness patterns. In this paper, we present $k$-POD, a novel method of $k$-means clustering on partially observed data that employs a majorization-minimization algorithm to identify a clustering that is consistent with the observed data. By bypassing the completely observed data formulation, $k$-POD retains all information in the data and avoids committing to distributional assumptions on the missingness patterns. Our numerical experiments demonstrate that while current approaches yield good results at low levels of missingness, some fail to produce results at larger overall missingness percentages or require prohibitively long computation time. By contrast, $k$-POD is simple, quick, and works even when the missingness mechanism is unknown, when external information is unavailable, and when there is significant missingness in the data.