An Empirical Evaluation of Techniques for Feature Selection with Cost

Stephen Adams, Ryan Meekins, Peter A. Beling · 2017

Feature selection is the process of selecting a subset of relevant features from the larger set of collected features. As the amount of available data grows with technology, feature selection becomes a more important part of the system-design process. In real-world applications, there are several costs associated with the collection, processing, and storage of data. Given that these costs can vary between data streams, it is important to consider the cost of features when performing feature selection. A majority of the feature selection algorithms select a relevant feature subset solely based on the merit and do not consider cost. In this study, we evaluate a previously proposed cost-based feature selection framework. We expand on the previously conducted experiments by testing a wider range of feature selection methods paired with the cost-based framework, testing a variety of classifiers, and sequentially adding features to the relevant subset based on the results of the cost-based framework. We find that the selection of the weight parameter that balances the effect of feature merit versus cost is tied to the choice of feature selection technique. The weight must be appropriately scaled with the value of the merit. Further, we confirm the previously tested results and offer insight into future research directions on the topic of feature selection and cost.

Read the paper · More papers on PaperTik