Clustering Random Variables

Richard J. Hathaway · IETE Journal of Research · 1998

The fuzzy c-means (FCM) clustering algorithm has long been used to cluster numerical data. More recently FCM has found application in certain schemes for clustering heterogenous data sets consisting of mixtures of numerical, interval, and fuzzy data. The range of applicability of FCM is extended here to include clustering data whose features are continuous random variables. Parametric, nonparametric and empirical models are presented, and in each case, the distributional laws of the random variables are encoded to give a real, finite-dimensional representation to which FCM can be applied. Subsequent decoding of the results yields cluster prototypes that are similar in nature to the original data themselves. Some properties of the approach are noted and the results of preliminary computational experimentation are given.

Read the paper · More papers on PaperTik