Some remarks on theR2for clustering
Nicola Loperfido, Thaddeus Tarpey · Statistical Analysis and Data Mining The ASA Data Science Journal · 2018
A common descriptive statistic in cluster analysis is theR2that measures the overall proportion of variance explained by the cluster means. This note highlights properties of theR2for clustering. In particular, we show that generally theR2can be artificially inflated by linearly transforming the data by “stretching” and by projecting. Also, theR2for clustering will often be a poor measure of clustering quality in high‐dimensional settings. We also investigate theR2for clustering for misspecified models. Several simulation illustrations are provided highlighting weaknesses in the clusteringR2, especially in high‐dimensional settings. A functional data example is given showing how thatR2for clustering can vary dramatically depending on how the curves are estimated.