Improving the co-training algorithm to enhance semi-supervised learning results
Ragini Kihlman, Maria Fasli · 2022 IEEE International Conference on Big Data (Big Data) · 2022
The co-training algorithm is one of the most common methods of semi-supervised learning in machine learning, which allows multiple learners to collaborate to discover the best information in unlabelled data. Co-training works well if the two views satisfy the sufficiency and independence assumptions. As a result of these assumptions of co-training and advancements in data classification algorithms, the performance of the underlying model could be improved. Specifically, view division, correlation between features in each view, domain knowledge, and label confidence estimation are introduced as key steps in improving co-training algorithms in this paper. Furthermore, we discuss the problems with the co-training methods currently being used, suggest some improvements, and speculate at how the algorithm could be improved going forward.