Comparative analysis of Gaussian mixture model, logistic regression and random forest for big data classification using map reduce

Vikas Singh, Rahul Kumar Gupta, Rahul K. Sevakula, Nishchal Kumar Verma · 2016

In the era of modern world, big data becomes major transformation of new technology, The amount of data generated by mankind is growing every year. To classify such big data is a challenging task with standard data mining techniques. This paper presents a Map Reduce based algorithm with Gaussian mixture model (GMM), Logistic regression(LR) and Random forest classifier (RFC). While, map phase determines the probabilities and class labels of the test data, the reduce phase predicts the class labels of test data by aggregating results from all the mappers. We have analyzed these algorithms on the basis of test accuracy, run time and number of mappers on multiple big data sets.

Read the paper · More papers on PaperTik