Learning from Small-sized Multi-class Positive Data and Unlabeled Data

Cheong-Hee Park · Journal of Korea Multimedia Society · 2023

To model a high performance classifier, it is necessary to have a sufficient amount of data belonging to each class. However, it takes a lot of time and effort to label all collected data with class labels. When labeled multi-class positive data and a lot of unlabeled data are given, it is common to assume that the unlabeled data contains negative data samples that do not belong to any positive class. In this paper, we propose a multi-class positive and unlabeled learning (MPUL) method in high dimensional data to learn a classifier that predicts positive and negative classes from small-sized multi-class positive data and large amounts of unlabeled data. The proposed method performs the expansion of the labeled data by selecting reliable positive and negative data from the unlabeled data and it improves the classification performance alleviating the problem of imbalance between classes. Experimental results using text data showed that the proposed method has obtained high performance improvement compared to the other methods.

Read the paper · More papers on PaperTik