An efficient algorithm for induction with random sampling
Ali Mirza Mahmood, Mrithyumjaya Rao Kuppa, V Sai Phani Chandu · 2011
Decision trees induction is among powerful and commonly encountered architecture for extracting of classification knowledge from datasets of labeled instances. However, learning decision trees from large irrelevant datasets is quite different from learning small and moderate sized datasets. In this paper, we propose a simple yet effective composite splitting criterion equal to a random sampling approach and gain ratio. Our random sampling method depends on small random subset of attributes and it is computationally cheap to act on such a set in a reasonable time. The superiority of the composite splitting criterion can persist when used for high dimensional datasets with irrelevant attributes. The empirical and theoretical prospective are validated by using 40 UCI datasets. The experimental results indicate that the proposed new heuristic function can result in much more simpler trees with almost unaffected or improved accuracy.