A Co-Training-based Algorithm Using Confidence Values to Select Instances
Karliane Medeiros Ovidio Vale, Flavius L. Gorgônio, Yago N. Araujo, Arthur C. Gorgônio, Anne M. P. Canuto · 2020
Data classification tasks have been used in a wide range of problems in the Machine Learning field and the different learning paradigms (supervised, unsupervised or semi-supervised) define the task a computer can learn from a set of labeled and/or unlabeled data. This paper presents a study in the semi-supervised learning paradigm and proposes changes on the co-training algorithm in order to propose a confidence value procedure to include new instances in the labeled dataset. In order to evaluate the proposed method, an empirical analysis with 30 datasets has been conducted, with different characteristics, that were set up with different percentages of initially labeled instances. Each dataset was trained using four different classification algorithms (Naive Bayes, Decision tree, Ripper and k-NN) as basis for the co-training training procedure. The obtained results are promising and they indicate that, in most cases, the proposed method performs better than the co-training method originally proposed in the literature.