Instance-based learning techniques of unsupervised feature weighting do not perform so badly!

Héctor Núñez, Miquel Sànchez–Marrè · 2004

Abstract. The major hypothesis that we will be prove in this paper is that unsupervised learning techniques of feature weighting are not significantly worse than supervised methods, as is commonly believed in the machine learning community. This paper tests the power of unsupervised feature weighting techniques for predictive tasks within several domains. The paper analyses several unsupervised and supervised feature weighting techniques, and proposes new unsupervised feature weighting techniques. Two unsupervised entropy-based weighting algorithms are proposed and tested against all other techniques. The techniques are evaluated in terms of predictive accuracy on unseen instances, measured by a ten-fold cross-validation process. The testing has been done using thirty-four data sets from the UCI Machine Learning Database Repository and other sources. Unsupervised weighting methods assign weights to attributes without any knowledge about class labels, so this task is considerably more difficult. It has commonly been assumed that unsupervised methods would have a substantially worse performance than supervised ones, as they do not use any domain knowledge to bias the process. The major result of the study is that unsupervised methods really are not so bad. Moreover, one of the new unsupervised learning method proposals has shown a promising behaviour when faced against domains with many irrelevant features, reaching similar performance as some of the supervised methods. 1

Read the paper · More papers on PaperTik