A novel parallel implementation of Naive Bayesian classifier for Big Data

Vijay D. Katkar, Siddhant Vijay Kulkarni · 2013

Big Data has become one of the most commonly used terms in the Information Technology circles. The sheer volume of data to be processed and analyzed has grown exponentially with the increasing popularity of Internet and World Wide Web. This presents challenges while storing, manipulating and mining the data. Every day researchers are working on different solutions to handle the volume of data being provided. In machine learning, classification of new observations is done on the basis of the provided learning(training) data to the classifiers. One of the most commonly used efficient and accurate classifiers is the Naive Bayesian classifier. This paper proposes a novel parallel implementation of Naive Bayesian (PNB+) classifier to decrease the testing time complexity while handling large data sets.

Read the paper · More papers on PaperTik