A problem of selecting optimal subset of fuzzy-valued features
Xiangkun Wang, Edward Tsang, D.S. Yeung · 2003
Feature subset selection refers to a data mining enhancement technique which aims to reduce the number of features to be used. This reduction is expected to improve the performance of data mining algorithms to be used, in aspects of speed, accuracy and simplicity. Although there has been some work on feature subset selection, research into the theoretically computational complexity of this problem and on the optimal selection of fuzzy-valued feature subsets has not been carried out. This paper focuses on a problem called optimal fuzzy-valued feature subset selection (OFFSS) which is regarded as being important but difficult in machine learning and pattern recognition. The measure of the quality of a set of features is defined by the overall overlapping degree between two classes of examples and the size of feature subset. The main contributions of this paper are that: (1) the concept of fuzzy extension matrix is introduced; (2) the computational complexity of OFFSS is proved to be NP-hard; (3) a simple but powerful heuristic algorithm for OFFSS is given; and (4) the feasibility and simplicity of the proposed algorithm are demonstrated via applications of OFFSS to input selection of neuro-fuzzy systems and to fuzzy decision tree induction.