Study of Influence of Dimension Reduction of High Dimensional Datasets in Classification Problem
Mohd. Salman Hossain Bhuiyan, Nabil Al Raian, Shahad Iqbal Leon, Musharrat Khan · 2020
We often develop many applications based on datasets. In case of high dimensional dataset, we face some problems while building any data mining model. When there are too many attributes in the dataset, then there may be dependency between attributes. There may be some irrelevant attributes too. Because of the influence of the dependent and irrelevant attributes, we get less accuracy in the data mining process. To overcome this problem, we need to reduce the dimensions of the dataset. In this work, we experimentally tested influence of dimension reduction on classification problems. For this purpose, we used four different datasets. We used backward elimination method to reduce the dimension of the dataset down to seven dimensions. We have experimented with Multi-layer Perceptron, Naïve Bayes, Decision Tree, Random Forest, K-Nearest Neighbor, and Support Vector Machine classification methods. We used 10-fold validation to train and test our dataset. Experimental results show that when the dimension is reduced, then the accuracy is improved for some classification algorithms like Multi-layer Perceptron, Naïve Bayes, and Random Forest. Our general finding is that if we exclude the less significant and irrelevant attributes, then the classification model gives better accuracy than it does without dimension reduction.