A MapReduce based approach for classification
Akash Neelesh Haldankar, Kiran Bhowmick · 2016
Classification is the process of developing a model which assigns records in a collection to predefined categories or classes. Various algorithms like K Nearest Neighbour, Naive Bayes, Support Vector Machine, and C4.5 have been implemented to develop a classification model. The ultimate goal of a classifier is to accurately predict the target class for a given set of input data. The process of classification begins with the creation of a model by taking into consideration an input as a training dataset. MapReduce is a programming model developed to process large datasets on distributed clusters of commodity machines. This paradigm can be used to implement a classifier. In this paper we study classification algorithms based on MapReduce and later introduce a MapReduce based Naive Bayes classification technique.