Hierarchical classification for dealing with the Class imbalance problem

Mohamed Bahy Bader-El-Den, Eleman Teitei, Mo Adda · 2016

The aim of classification in machine learning is to utilize knowledge gained from applying learning algorithms on a given data so as determine what class an unlabelled data having same pattern belongs to. However, algorithms do not learn properly when a massive difference in size between data classes exist. This classification problem exists in many real world application domains and has been a popular area of focus by machine learning and data mining researchers. The class imbalance problem is further made complex with the presence of associative data difficult factors. The duo have proven to greatly deteriorate classification performance. This paper introduces a two-phased data level approach for binary classes which entails the temporary re-labelling of classes. The proposed approach takes advantage of the local neighbourhood of the minority instances to identify and treat difficult examples belonging to both classes. Its outcome was satisfactory when compared against various data-level methods using datasets extracted from KEEL and UCI datasets repository.

Read the paper · More papers on PaperTik