A survey on state of art approaches in handling imbalance, positive and unlabelled data
N. Deepa, R. Sumathi · 2022
Machine learning algorithm extends its horizon in various application domains. Real time applications in various sectors which can be effectively handled with machine learning techniques lack from it due to the unavailability of the required data. In such cases the data may be imbalanced, with a single class labeled and other unlabelled. However, this will be the prevailing scenario in future and has to be addressed. This work concentrates on studying the various approaches that are in practice to handle these kinds of data. In specific, it concentrates on approaches used for handling imbalance data sets and the recent works in it, methods used for handling positive unlabelled data. Though there are few solutions in practice for handling imbalance data and positive unlabelled data. Approaches that work at the data level and algorithm level approach for handling the former and techniques such as biased learning and class prior models for handling the later are in practice. One class classifier is another approach that can be used for addressing the problem that arises with the class imbalance in the dataset and related issues. It has also been noted from the literature that one class classifier will be more suitable but it has not been widely used in addressing the problems such as class rarity. It can be inferred from this study that the one class classifier would play a significant role in various domains in future.