A Simple Filter Benchmark for Feature Selection
Athanasios Tsanas, Max A. Little, Patrick McSharry · 2010
A new correlation-based filter approach for simple, fast, and effective feature selection (FS) is proposed. The association strength between each feature and the response variable (relevance) and between pairs of features (redundancy) is quantified via a simple nonlinear transformation of correlation coefficients inspired by information theoretic concepts. Furthermore, the association strength between a set of features and the response variable (feature complementarity) is explicitly addressed using a similar nonlinear transformation of partial correlation coefficients, where a feature is selected conditionally upon its additional information content when combined with the features already selected in the forward sequential process. The new filter scheme overcomes several major issues associated with competing FS algorithms, including computational complexity and difficulty in implementation, and can be used on both multi-class classification and regression problems. Experiments on five synthetic and twelve real datasets demonstrate that the proposed filter outperforms popular alternative filter approaches in terms of recovering the correct features. We envisage the proposed scheme setting a competitive benchmark against which more sophisticated FS algorithms can be compared. Documented Matlab source code is available on the first author’s website.