Treating missing data processing based on neural network and AdaBoost
Miao Zhimin, Pan Zhisong, Guyu Hu, Zhao Lu-wen · 2007
Missing data is a common problem in data quality. Such data are generally ignored or simply substituted in classification problem, which will affect the performance of a classifier. In the paper an innovative framework RBP- AdaBoost for handling with missing features values in classification is presented. This framework is composed of two parts: predicting the missing values and classifying the data including predicted missing values. Back-propagation algorithm (BP) is adopted to predict missing value firstly, and Adaptive Boosting (AdaBoost) as a methodology of aggregation of many weak classifiers into one strong classifier is used in classifying predicted missing data. We carry out experiments with nine UCI datasets to evaluate the effect on classification error rate of four general methods and the prediction model of BP. Experimental results show that the classification rate of the proposed new framework RBP- AdaBoost is increased 6.4% to 23.69% comparing with other methods. The performance of missing data treatment model is considered to be effective.