Bayesian classification using an entropy prior on mixture models
Julian L. Center · AIP conference proceedings · 2001
In many classification problems, it is reasonable to base the analysis on a mixture model. A mixture model assumes that each sample is produced by first randomly selecting from a finite collection of data clusters and by then using the chosen cluster distribution to produce the class label and feature vector of the sample. If we know the set of model parameters, then when we observe a feature vector, we can predict the classification. When we do not know the parameters exactly, we must infer the model parameters from a training set of data samples. Taking the Bayesian approach, we want to determine the probability distribution for the parameters given the training data. Then when it comes time to predict the class label, given a feature vector, we integrate over the model parameter distribution. We argue that a good, objective choice for the prior distribution on the model parameters is based on the entropy of each mixture model. We show that this prior regularizes the model fit so that over-fitting the training data has no adverse effects.