A scalable random forest algorithm based on MapReduce

Jiawei Han, Yanheng Liu, Xin Sun · 2013

Random Forest is a popular data classification algorithm for machine learning. This paper proposes SMRF algorithm--an improved scalable Random Forest algorithm based on Map Reduce model. This new algorithm makes data classification in computer cluster or cloud computing environment for massive datasets. SMRF processes and optimizes the subsets of the data across multiple participating computing nodes by distributing. The experimental results show that the SMRF algorithm has the equally accuracy degradation but higher performance while comparing with traditional Random Forest algorithm. SMRF algorithm is more suitable to classify massive data sets in distributing computing environment than traditional Random Forest algorithm.

Read the paper · More papers on PaperTik