Classification in a Distributed System - A Study of Random Forest in the Hadoop MapReduce Framework
K.N.P. Kumar, Neeraj Anand Sharma, Aamir Ali · 2019
Due to the massive increase in the amount of data collected and stored globally, new issues of data processing and algorithm compatibility has come to light. Processing large datasets is a challenge for commodity systems due to its limited processing capabilities while specialized and more powerful processing systems are costly to set up and use. This research will focus on a distributed system known as the Hadoop Framework that uses commodity machines to create a combined and more powerful system that is able to process big datasets much more efficiently. The Hadoop Framework processes data in a distributed and parallel environment which is unlike the processing mechanisms used in a traditional commodity computer system. However, this change in the nature of processing demands a change in the underlying structure of the machine learning algorithms that process the data. The research will delve at implementing the Random Forest algorithm in the Hadoop Framework. The Random Forest algorithm will be executed on a four-node Hadoop Cluster inputting data of varying sizes at a time. Performance of the algorithm will be measured by analyzing execution Time, Accuracy, Kappa, Reliability and Standard Deviation measures. Results will show that the execution time reduces constantly as the size of dataset increases thus making the system a key tool for processing large datasets.