STDPboost: A Self-Training Method Based on Density Peaks and Improved Adaboost for Semi-Supervised Classification
Lin Xu, Junnan Li · IEEE Access · 2023
The self-training methods have been praised by extensive research in semi-supervised classification. Mislabeling is the main challenge in self-training methods. Multiple variations of self-training methods are recently proposed against mislabeling from the following one of two aspects: a) using heuristic rules to find high-confidence unlabeled samples that can easily be predicted correctly in each iteration; b) enhancing prediction performance by employing ensemble classifiers composed of multiple weak classifiers. Yet, they still suffer from the following issues: a); most strategies for finding high-confidence unlabeled samples heavily rely on parameters; b) almost all employed ensemble classifiers originally designed for supervised classifiers and may not be suitable for semi-supervised classification due to the limited number and unrepresented distribution of the initial labeled data; c) few can overcome mislabeling from the above two aspects at the same time. To advance the state of the art, a new self-training method based on density peaks clustering and improved Adaboost is presented and named as STDPboost. In the iterative self-taught process, a new density peaks clustering-based strategy is proposed to find high-confidence unlabeled samples and a new ensemble classifier named AdaboostSEMI and more suitable for semi-supervised classification is proposed to predict high-confidence unlabeled samples, which overcomes mislabeling and the mentioned shortcomings of existing self-training methods. Intensive experiments on benchmark data sets have proven that STDPboost outperforms 7 state-of-the-art self-training methods in average classification accuracy of KNN classifier and CART classifier with the percentages of the initial labeled data from 10% to 50% due to further alleviating mislabeling.