Detecting Low Complexity Clusters by Skewness and Kurtosis in Data Stream Clustering.
Mingzhou Joe Song, Hongbin Wang · ISAIM · 2006
Established statistical representations of data clusters employ up to second order statistics including mean, variance, and covariance. Strategies for merging clusters have been largely based on intraand inter-cluster distance measures. The distance concept allows an intuitive interpretation, but it is not designed to merge from the viewpoint of probability distributions. We suggest an alternative strategy to compare clusters based on higher order statistics to capture the underlying probability distributions. Higher order statistics, such as multivariate skewness and kurtosis, enable a more accurate description of the shape of a cluster. Although the original definitions of kurtosis and skewness do require simultaneous involvement of all data points, our finding shows that their estimation can be decomposed into combinations of the cross moments of subsets of data. This decomposable property makes it possible to apply skewness and kurtosis to data stream clustering, where historical data are not accessible. We utilize tests for normality based on skewness and kurtosis to discover cluster pairs that can be merged to produce a less complex normal cluster even if they have different means or covariance structures.