Semisupervised Learning of Mixture Models with Class Constraints
Qi Zhao, David J. Miller · 2006
Most prior work on semisupervised clustering/mixture modeling with given class constraints assumes the number of classes is known, with each learned cluster assumed to be a class and, hence, subject to the given instance-level constraints. When the number of classes is incorrectly assumed and/or when the "one-cluster-per-class" assumption is not valid, the use of constraint information in these methods may actually be deleterious to learning the ground-truth data groups. We extend semisupervised learning with constraints (1) to allow allocation of multiple mixture components to individual classes and (2) to estimate both the number of components/clusters and, leveraging the constraint information, the number of classes present in the data. For several real-world data sets, our method is shown to estimate correctly the number of classes and to give a favorable comparison with the recent mixture modeling approach of N. Shental et al. (see NIPS, 2003).