Feature Selection Using Heterogeneous Data Indexes: a data science perspective
Divyansh Saxsena, Mithileysh Sathiyanarayanan · 2019
The curse of dimensionality is the problem of keeping too many attributes in a dataset with a number of instances from data science (analytics) perspective. Most of the attributes (or features) may not be significant or even can harm the process of classification by orienting this process to a wrong direction and subsequently yielding poor classification accuracy. Therefore, it is essential to retain only meaningful features in a dataset. In this paper, we propose a novel approach for selecting a subset of features from a dataset which is supervised in nature, having labels for each of its pattern. Five methods including data indexes viz. Information Gain, Laplace Score, Variance, Cosine Similarity and Fisher Score (one index in one method) are applied to find scores and hence ranks of all features in a dataset. The set of features (with ranks arranged in descending orders) is divided in four feature subsets with size of maximum 100,75,50 and 25% respectively. Apply statistical mode approach to fill each subset from the full subset of features. Thus, we get four subsets of features. Each of the four subsets are tested using a k - nearest neighbor classifier. The results in terms of classification accuracies and processing time for all the four subsets are obtained. After extensive experiments of proposed approach on nine datasets, the encouraging results justify the significance and suggest for further implementation on datasets with high dimensionality.