Algorithm of decision trees insensitive to data distribution
Lijuan Liu · Journal of Jilin University · 2009
Traditional decision tree algorithms are sensitive to data distribution.The predictive accuracy of minority class is often decreased when the algorithm deals with skewed datasets.There exist some algorithms which can only handle the skewed datasets with only two kinds of classes.A new decision tree algorithm called DTID is proposed,which is insensitive to data distribution.Using this algorithm new cases of each minority class are generated to adjust the data distribution of the sample set,and the predictive accuracy of each minority class is improved.By adopting the modulus of each case previously,the running time of the algorithm is reduced.Experimental results show that,compared with C4.5 algorithm,the accuracy of DTID is obviously improved and it can obtain much better result even though there are many minority classes in the sample set.