An analytical approach to similarity measure selection for self-training
Vincent Van Asch, Walter M. P. Daelemans · 2013
We present a framework for investigating properties of similarity measures as a criterion for selecting the best-suited measure for a specic task, in this paper: corpus selection for self-training. We focus on the squared Pearson’s correlation coecient as the property to rank similarity measures. Selftraining is an unsupervised domain adaptation technique, in which three corpora are involved. Especially, the choice of the unlabeled corpus can be important and we show that similarity measures can be helpful when selecting an unlabeled corpus. In addition, we found that the correlation coecient between similarity and accuracy of a similarity measure can be used to select the most suitable similarity measure, but other properties of similarity measures do also play a role.