Unsupervised learning of Gamma mixture models using Minimum Message Length

Yudi Agusta, David L. Dowe · 2003

Mixture modelling or unsupervised classification is a problem of identifying and modelling components in a body of data. Earlier work in mixture modelling using Minimum Message Length (MML) includes the multinomial and Gaussian distributions (Wallace and Boulton, 1968), the von Mises circular and Poisson distributions (Wallace and Dowe, 1994, 2000) and the distribution (Agusta and Dowe, 2002a, 2002b). In this paper, we extend this research by considering MML mixture modelling using the Gamma distribution. The point estimation of the distribution was performed using the MML approximation proposed by Wallace and Freeman (1987) and gives impressive results compared to Maximum Likelihood (ML). We then considered mixture modelling on artificially generated datasets and compared the results with two other criteria, AIC and BIC. In terms of the resulting number of components, the results were again impressive. Application to the Heming Pike dataset was then examined and the results were compared in terms of the probability bitcostings, showing that the proposed MML method performs better than AIC and BIC. A further application also shows that our method works well with datasets containing left-skewed components such as the Palm Valley (Australia) image dataset.

Read the paper · More papers on PaperTik