Random forest in semi-supervised learning (Co-Forest)
Nesma Settouti, Mostafa El Habib Daho, Mohammed El Amine Lazouni, Mohammed Amine Chikh · 2013
The semi-supervised learning has been widely applied in many fields such as medical diagnosis, pattern recognition. The semi supervised learning methods are used to employ unlabelled data in addition to labelled data for better classification of large data sets, where only a small number of labelled examples is available. Ensemble Methods are considered as an effective solution to the problem of dimensionality and can improve the robustness and generalization ability of individual learners. In this paper, we are particularly interested in the overall algorithm Random Forest semi-supervised named Co-Forest for the classification of large biological data. The algorithm is evaluated on its ability to correctly predict the labels of unlabelled examples, and its robustness when the number of labelled examples available decreases.