Semi-Supervised Learning of Mixture Models and Bayesian Networks
Fábio Gagliardi Cozman, Ira L. Cohen, Marcelo César Cirelo · 2003
This paper analyzes the performance of semisupervised learning of mixture models. We show that unlabeled data can lead to an increase in classification error even in situations where additional labeled data would decrease classification error. This behavior contradicts several empirical results reported in the literature. We present a mathematical analysis of this "degradation" phenomenon and show that it is due to the fact that bias may be adversely affected by unlabeled data.