A Distributed Decision Tree on Hadoop

William Jones, Min Chen · 2021

Decision trees have become a very popular form of supervised learning due to their ease of implementation and understandability, ability to work with non-normalize data, and low runtime when predicting class labels. Unfortunately, these benefits come at the cost of a large runtime when constructing the tree that greatly increases as the size of the data set increases. The objective of this study is to implement a decision tree using Hadoop, which can greatly reduce runtime over large data sets and offers scalability for future changes in the size of the data.

Read the paper · More papers on PaperTik