Separating the wheat from the chaff

Sean Stijven, Wouter Minnebo, Katya Vladislavleva · 2011

Feature selection in high-dimensional data sets is an open problem with no universal satisfactory method available. In this paper we discuss the requirements for such a method with respect to the various aspects of feature importance and explore them using regression random forests and symbolic regression. We study 'conventional' feature selection with both methods on several test problems and a case study, compare the results, and identify the conceptual differences in generated feature importances.

Read the paper · More papers on PaperTik