A Map-Reduce Model of Decision Tree Classifier using Attribute Partitioning
N Shivaraju, Vijayakumar Kadappa, Shankru Guggari · 2017 International Conference on Current Trends in Computer, Electrical, Electronics and Communication (CTCEEC) · 2017
Data mining is a process of analyzing data to extract the patterns in large datasets in the field of artificial intelligence, machine learning and statistics. Decision tree is one of the well-established classification models in data mining. The size and dimensionality of the data of today's world are increasing exponentially, thus finding of informative patterns is an important and crucial task. The organizations require distributed systems for storing and processing huge amount of data. The proposed method is the parallel implementation of Decision Tree methods based on the idea of attribute partitioning, where dataset is partitioned into multiple subsets of dimensions. We develop a Map-Reduce programming model for processing data using decision tree classifier based on attribute partitioning. The experimental results show an improvement in classification accuracy as compared to a traditional decision tree method.