Model complexity validation for PDF estimation using Gaussian mixtures
L. Sardo, Josef Kittler · 2002
Semiparametric density estimation using Gaussian mixtures is a powerful means that can give as good performance as a nonparametric estimator, without its heavy computational burden. A maximum penalised likelihood principle was previously proposed by the authors (1996) for selecting the best approximating mixture for an unknown density function. We propose here a test carried on the training set to validate the model choice. The selected model is required to give a calibrated prediction, i.e. if it predicts the frequencies of the training sample reasonably well, the penalty term adopted is accepted otherwise it is relaxed.