Ensemble methods in multi-label classification
A. Hermann Müller · SUNScholar (Stellenbosch University) · 2018
There are many scenarios where several labels may be associated simultaneously with each data case in a dataset.Therefore, a large number of multi-label datasets are found in a variety of domains, including image annotation, text annotation, bioacoustics, music research and medical diagnostics.In this thesis, we focus on such multi-label datasets with the particular goal of performing multi-label classification.Multi-label classification is an extension of binary-and multiclass classification to scenarios where each instance in a dataset can have multiple or none of K labels.Methods to perform multi-label classification can be divided into three categories: problem transformation methods, algorithm adaptation methods and multi-label ensemble methods.Multilabel ensemble methods receive particular attention in this research.We discuss previously proposed multi-label ensemble methods and also propose a new multi-label ensemble method, named label dependent splitting (LDsplit) with trees.LDsplit with trees constructs an ensemble of tree-structures by considering different permutations of the labels in a multi-label dataset.The method differs from other multi-label ensemble methods since each tree-structure splits the multi-label data in a label-dependent way, whilst incorporating label correlation.Furthermore, each split of a node in a tree-structure is performed by considering a binary classification problem.By performing an empirical study on benchmark datasets, the predictive performance of LDsplit with trees is compared to that of other multi-label learning methods.LDsplit with trees produces very promising results, allowing us to believe that with further modifications the procedure may become a highly competitive multi-label learning method.Furthermore, we also explore aspects of analysing text data.We perform an extensive analysis on a practical multi-label text dataset.The practical dataset consists of online comments, where each comment is labeled to identify if any or multiple so-called "toxicity" are present in the comment.Our model may therefore be used to identify different types of toxicity present in online comments and help reduce online abuse and harassment.A particular challenge faced in the practical data analysis is the sparsity of the labels.