Comparative analysis of attribute selection measures used for attribute selection in decision tree induction
A. S. Bhatt · 2012
Data mining is a process of finding hidden information from databases storing historical data which are also known as data-warehouses. Classification being a very well-known data mining technique, groups similar data objects by establishing relationship between the objects under test and the pre-defined class labels obtained during training phase. Of all the classification algorithms, decision tree is most commonly used. In this paper we will discuss scalability of decision tree algorithm based on the selection of the attribute selection measure. Attribute selection measure is mainly used to select the splitting criterion that best separates the given data partition. The popular attribute selection measures are Information Gain and Gain Ratio. We would perform the comparative analysis of these measures and based on their results we would determine which measure should be used in which situation in order to increase the scalability of the Decision Tree algorithm.