Imbalanced Classification Problem Using Data-driven and Random Forest Method
Wan Wang, Xinglu Liu, Wai Kin Victor Chan · 2020
Classification problem is a major concern in numerous domains, for instance, customer retaining problem on electronic devices, rare disease diagnosis, bank fraud identification etc. Accurate and efficient classification approaches can reduce operation cost of companies and thus save a mass of artificial work significantly. Identifying the potential targets and then put promotion strategy to customers in order to achieve considerable improvement has become a critical objective in a range of institutions. Nevertheless, missing values and some statistical errors are always involved in the majority of datasets. Besides, the available datasets usually tend to be imbalanced, which leads to weak performance of the classification model. Though random forest method is widely adopted in prediction problem, few of them introduce data-driven technique in this problem. In th is work, firstly, a data-driven method is applied to generate new features based on the imbalanced datasets. Then, the prediction is executed by Random Forest(RF) algorithm on three open-sourced datasets. Results reveal that random forest combined with data-driven method improved the prediction accuracy. The AUC values also perform well stable and even increased. Our main contributions are: We consider two-class classification problem and employ a data-driven method to generate new features, which can raise the correlation of the features and increase the accuracy without losing original information meanwhile. 2. We use random forest coupled with a data-driven method, and consider imbalanced condition in our test which are more reasonable in real world applications, which can provide a general idea and direction in the industry.