Neighborhood graph construction for semi-supervised learning
Lilian Berton, Alneu de Andrade Lopes · AI Matters · 2016
Semi-supervised learning (SSL) is useful when few labeled and plenty of unlabeled examples are available. This occurs in most of the cases due to labeled instances be difficult, expensive and time consuming to be obtained since human experts are required for the labeling task (Chapelle, Schlkopf, & Zien, 2010). The training set contain labeled data represented by X L = {( x 1 , y 1 ) ... ( x 1 , y 1 )}, and unlabelled data represented as X U = { x l+1 ... x l+u }. The total amount of training data is X = l ∪ u . The set of labeled examples is associated with the labels L Y = { y 1 , ..., y l } where y i ε {1,..., c } and c is the number of classes. The purpose of SSL is to infer the missing labels Y U = { y l+1 ,..., y n } corresponding to the unlabelled set X U .