Prior Dirichlet Distribution Based Feature Selection in Naive Bayes on High Dimensional Data Classification
Hui Chen, Lifei Chen, Qingshan Jiang, Shun Guo · 2020
Classification is an essential technology in data mining, which has been widely used in real-world problems. A massive amount of information is pouring into our lives since the development of science and technology. The data size and dimension are growing rapidly in various fields and challenging data mining and machine learning approaches. Feature selection reduces the dimension of data and selects significant features, which is an important research direction. However, some norm measures of feature datasets may be imprecise. Many feature selection methods have the characteristics of high cardinality and high computational complexity. Therefore, there is a strong need to study effective feature selection methods. Naive Bayes is a simple and effective learning theory that does not need various parameters. However, its conditional independence assumption limits its classification results in practical applications. Meanwhile, the curse-of-dimensionality brings noises and messy interaction between features. In this paper, we propose a feature selection method named as the Prior Dirichlet Distribution of Naive Bayes (PDDNB) which optimizes the Bayesian theory based on Dirichlet priori distribution. This method simplifies the computational ways to an analytical solution, which significantly improve the calculation efficiency. The executive experiments on the ALL-AML datasets and the other five real-world datasets turn out our method to effectively increase classification accuracy.