Adaptive Naive Bayesian Classifier for Automatic Classification of Webpage from Massive Network Data
Xu LinBin, Jun Liu, Zhou WenLi, Qing You Yan · 2014
This paper presents the application of Naïve Bayesian classifier to automatic classification of webpage. The key point in this article is that massive empirical data derives from the real traffic data collected from the backbone network of certain province in China, and we apply cumulative probability to determine the optimal size of feature vector adaptively. It's proved that the adaptive method of cumulative probability threshold selection applied in this study has good robustness. This paper focus on four feature selection methods: TF-IDF (term frequency-inverse document frequency), IG (Information Gain), MOR (Multi-class Odds Ratio), CDM (Class Discriminating Measure). We find that Naïve Bayesian classifier performs fairly well in speed and precision on big data sets, whose precision, recall and F1 metric are all above 90% in all 6 categories of webpage.