An Optimized Classification and Regression Tree Algorithm by Combining Feature Selection Methods
Noha Khamis Elsayed Yakout, Magda Mohamed Madbouly, Mohamed El Sherbiny · Journal of Computing and Communication · 2025
Feature selection is the process of removing features from the data set that are irrelevant with respect to the task that is to be performed. Also, Feature selection can be extremely useful in reducing the dimensionality of the data to be processed by the classifier, reducing execution time and improving predictive accuracy . In addition, Feature selection is a dimensionality reduction technique that reduces the number of attributes to a manageable size for processing and analysis . The accuracy of the classifier not only depends on the classification algorithm but also on the feature selection method. Selection of irrelevant and inappropriate features may confuse the classifier and lead to incorrect results. Feature selection is must in order to improve efficiency and accuracy of classifier . There are several classification methods. One of the famous methods of classification is decision tree. The decision tree is used for finding the best way to distinguish a class from another class. There are five mostly & commonly used algorithms for decision tree: - ID3, CART, CHAID, C4.5 algorithm and J48 . The CART (Classification And Regression Tree) is a nonparametric model which uses historical data to construct so-called decision trees. Trees are built top-down recursively beginning with a root node . In this paper, a new technique is suggested to optimize the classification of the classification tree of CART algorithm by using Combination of Feature selection methods which are Principal Component Analysis (PCA) method and Information Gain method. The PredictorImportance(imp) of decision tree express the accuracy of this tree. The proposed model is practiced on labor database. Results shows the classifier accuracy and predictor importance have been surely enhanced by the use of Feature selection methods than the classifier accuracy and predictor importance without feature selection.