SELECTION OF NUMERICAL AND NOMINAL FEATURES BASED ON PROBABILISTIC DEPENDENCE BETWEEN FEATURES

Krzysztof Michalak, Halina Kwaśnicka, Ewa Wątorek, Marian Klinger · Applied Artificial Intelligence · 2011

Data classification tasks often concern objects described by tens or even hundreds of features. Classification of such high-dimensional data is a difficult computational problem. Feature selection techniques help reduce the number of computations and improve classification accuracy. In Michalak and Kwasnicka (Citation2006a, Citationb) we proposed a feature selection strategy that selects features in an individual or pairwise manner based on the assessed level of dependence between features. In the case of numerical features, this level of dependence can be expressed numerically using linear correlation coefficients. In this paper, the feature selection problem is addressed in the case of a mixture of nominal and numerical features. The feature similarity measure used in this case is based on the probabilistic dependence between features. This similarity function is used in an iterative feature selection procedure, which we proposed for selecting features prior to classification. Experiments prove that using the probabilistic dependence similarity function along with the presented feature selection procedure can improve computation speed while preserving classification accuracy in the case of mixed nominal and numerical features.

Read the paper · More papers on PaperTik