Data Representation Using Mixtures of Principal Components
2009
This chapter presents a new approach to data representation, the mixture of principal components. In general, it can be considered as a portion of a spectrum of representations and has been successfully used in such diverse applications as image compression and nonlinear clustering. At one extreme of the spectrum is vector quantization (VQ). Data are represented by a set of zero-dimensional points, Voronoi centers, within the N -dimensional space. Only data corresponding to the exact values of the centers are represented exactly. As such, it is a nonlinear representation. Euclidean distance is used to measure similarity under this representation. At the other extreme lies principal component analysis (PCA). Here, data are represented by a linear combination of a set of N basis vectors. The representation is complete and continuous because all possible data vectors may be represented exactly. Between these two extremes lies the mixture of principal components (MPC). Data are represented by a set of M -dimensional subspaces where 0 M N . The subspace projection length is used as the similarity measure which reduces to the vector angle for the one-dimensional case. In this approach a data vector within a class (subspace) is represented as a continuous, linear combination of the M basis vectors of the subspace in a manner analogous to the PCA representation. But, because of the partitioning of the data into a discrete number of regions or classes, the MPC effects a nonlinear mapping of the data as does VQ. Applications presented include grayscale image feature extraction and color segmentation.