Hypothesis margin based weighting for feature selection using boosting
Malak Alshawabkeh · 2013
Feature selection (FS) is a preprocessing process aimed at identifying a small subset of highly predictive features out of a large set of raw input variables that are possibly irrelevant or redundant. It plays a fundamental role in the success of many learning tasks where high dimensionality arisesas a big challenge. Many endeavors to cope with this problem have been attempted and various outstanding feature selection methods have been proposed. Recently, there has been a growing line of research in utilizing the concept of hypothesis margins to measure the quality of a set of features. However, most previous feature selection algorithms have been developed under the large hypothesis margin principles of the 1-NN algorithm, such as Simba. Little attention has been paid so far to exploiting the hypothesis margins of boosting to evaluate features. Boosting is well known to maximize the training examples' hypothesis margins, in particular, the average margin that considers the whole margin distribution and thus include more information. In this thesis, we took an unusual approach for using boosting as an effective FS by utilizing the training examples' mean margins. A weight criterion, termed Margin Fraction (MF), is assigned to each feature that contributes to the margin distribution combined in the final output produced by boosting. We argue that using the MF is more favorable for several reasons. First, boosting hypothesis margins have been used both for theoretical generalization bounds and as guidelines for algorithm design, and thus, a natural goal is to find learners (features) that achieve a maximum margin. Second, current boosting-based feature selection methods measure the relative importance of features based on the Confidence Ratio (CR) of the learned base hypothesis. However, while a feature may have a large CR, it will not contribute to a good overall margin unless its "conditional" margin is also large.