Variable selection for regression problems using Gaussian mixture models to estimate mutual information
Emil Eirola, Amaury Lendasse, Juha Karhunen · 2014
Variable selection is a crucial part of building regression models, and is preferably done as a filtering method independently from the model training. Mutual information is a popular relevance criterion for this, but it is not trivial to estimate accurately from a limited amount of data. In this paper, a method is presented where a Gaussian mixture model is used to estimate the joint density of the input and output variables, and subsequently used to select the most relevant variables by maximising the mutual information which can be estimated using the model.