An Improved Naive Bayesian Classification Algorithm for Massive Data
Sun Tongjing, Ji Li, Ning Ke · 2018
For the low speed and accuracy in massive data classification, an improved Naive Bayesian classification algorithm for mass data processing is proposed. Firstly, feature rough clustering is carried out to cluster the features to reduce the computational complexity of feature association. Secondly, the association rules algorithm is used to mine frequent item sets of rough clustering subsets, and the generated frequent item sets are used to filter the features based on the result of classification. And then, the feature set after feature selection is weighted to improve the accuracy. Finally, the improved algorithm is implemented on the MapReduce parallelization platform and tested with five data sets of different sizes. The experimental results show that the improved algorithm in this paper could save a lot of running time when dealing with large-scale data sets, and maintain high accuracy.