A decision tree generation algorithm based on maximum similarity
Xinmeng Zhang, Shengyi Jiang · 2011
Node splitting is good or bad depends on the measure method of the impurity. We propose a new decision tree feature selection strategy based on maximum similarity, called fsms. First, splitting the dataset into subset according to each attribute value, calculating the sum of average similarity of each subset, then selecting the attribute with the maximum similarity as the best splitting attribute. Experimental results show, After tested in multiple test dataset, The decision tree constructed by the algorithm is better than Some classic algorithms such as id3,c4.5 in The classification precision, and less affected by the size of dataset.