Feature Extraction and Selection
Dick de Ridder, David M. J. Tax, Bangjun Lei, Guangzhu Xu, Ming Feng, Yaobin Zou, Ferdinand van der Heijden · 2017
The first step in the design of optimal feature selectors and feature extractors is to define a quantitative criterion that expresses how well such a selector or extractor performs. The second step is to do the actual optimization, which is to use that criterion to find the selector/extractor that performs best. This chapter introduces the problem of selecting a subset from the N-dimensional measurement vector such that this subset is most suitable for classification. It assumes that the class-dependent distributions are such that the expectation vectors of the different classes are discriminating. If fluctuations of the measurement vectors around these expectations are due to noise, then these fluctuations will not carry any class information. The chapter explores a search strategy called 'branch-and-bound' and continues with some of the much faster, but suboptimal, strategies. It explains the concept of performance measures. These performance measures evaluate the appropriateness of a measurement vector for the classification task.