Dealing with Imbalanceness in Hierarchical Multi-Label Datasets Using Multi-Label Resampling Techniques

Rodolfo Miranda Pereira, Yandre M. G. Costa, Carlos Nascimento Silla Junior · 2018

The task of learning from imbalanced datasets has been widely investigated in the binary, multi-class and multilabel scenarios. Although this problem also affects hierarchical datasets, to the best of our knowledge, there are no works in the literature that deal with imbalanceness in hierarchical contexts. In this paper we propose metrics to measure "how imbalanced" is a Hierarchical Multi-Label Dataset, in addition to an approach to deal with this imbalanceness using Multi-Label resampling techniques. The proposed technique is based on the conversion of the dataset labels to a strictly multi-label format, applying wellknown multi-label resampling techniques and then converting the dataset back to its hierarchical taxonomy. The experimental evaluation over a highly imbalanced Music Genre Recognition dataset achieved promising results, with an increase of 0.2337 in the Avg-AUROC metric in relation to the original dataset.

Read the paper · More papers on PaperTik