Machine Learning on imbalanced data in Credit Risk

Shiivong Birla, Kashish Kohli, Akash Dutta · 2016

In Machine Learning, we often encounter instances of imbalanced data which occur whenever there is an unequal representation in the classification categories. New found interest in Machine Learning has made its usage ubiquitous. Its applications encompass a wide plethora of scenarios ranging from Business and Banking to Bioinformatics and Psychology. These problems are often characterized by imbalanced data, the presence of which often leads to inaccurate predictive models, since the distribution of testing data may differ from that of training data while learning, leading to misclassification of the response variable. The primary focus of the paper is on Credit Risk which is defined as the probability of defaulting on the loan or credit acquired from a banking or financial institution. The base risk is that of the loss of primary principal and interest, disruption of cash flows and increased collection costs. Loan Default is an uncommon phenomena, henceforth we obtain the imbalanced data. We've adopted the approach of Logistic Regression and Classification and Regression Trees (CART) with techniques such as undersampling, Prior Probabilities, Loss Matrix and Matrix Weighing to deal with imbalanced data.

Read the paper · More papers on PaperTik