Bayesian classification learning framework based on bias–variance trade-off

文钧 张, 良孝 蒋, 欢 张, Chengyu Hu · Scientia Sinica Informationis · 2022

Due to its simplicity, efficiency, and efficacy, naive Bayes (NB) continues to be one of the top ten data mining algorithms. However, its attribute-conditional independence assumption rarely holds true in real-world applications. In order to alleviate the need for this assumption, scholars have proposed five types of improved approaches, including structure extension, attribute selection, attribute weighting, instance selection, and instance weighting. Although these existing improved approaches reduce the bias of the model to some extent, they also increase the variance of the model and thus limit the generalization of the model. The bias–variance trade-off is one of the core principles of machine learning, which requires a model to have low bias and variance at the same time. This paper is focused on how to introduce the bias–variance trade-off into Bayesian classification learning, obtain lower bias and variance at the same time, and improve the generalization of the model. Therefore, we first theoretically analyze the feasibility of introducing the bias–variance trade-off into Bayesian classification learning and determine the key factor that ensures feasibility. Then, we learn the posterior probability losses of the Bayesian classification models by constructing regression tasks to control the change in the key factor. Finally, we propose a Bayesian classification learning framework based on the bias–variance trade-off and re-implement NB and its various improved models under the proposed learning framework. The experimental results on a large number of classical UCI standard datasets show that under the proposed learning framework, the classification performance of existing various state-of-the-art Bayesian classification models is significantly better than their original performance.

Read the paper · More papers on PaperTik