Boosting for Vote Learning in High-Dimensional kNN Classification
Nenad Tomašev · 2014
Intrinsically high-dimensional data has recently been shown to exhibit substantial hubness in terms of skewness of the k-nearest neighbor occurrence frequency distribution. While some points arise as centers of influence and dominate most k-nearest neighbor sets, other points occur very rarely and barely affect the inferred models. Hubness has been shown to be highly detrimental to many learning tasks and several hubness-aware learning methods have recently been proposed. This paper extends the existing fuzzy neighbor occurrence models in order to enable cost-sensitive learning. We evaluate the extended implementations within the context of multi-class boosting, which is used to learn the appropriate neighbor votes during the re-weighting iterations. The proposed approach is evaluated on a series of high-dimensional datasets from various domains. The results demonstrate promising improvements of the proposed approach over the baselines.