Scorecard construction with unbalanced class sizes
Veronica Vinciotti, David J. Hand · Journal of the Iranian Statistical Society · 2003
A long-running issue in scorecard construction is how to handle dramatically unbalanced class sizes. This is important because, in many applications, the class sizes are very different. For example, it is common to find that 'bad' customers constitute less than 10% of the customer base and even more extreme situations often arise: Brause et al (1999) remark that in their database of credit card transactions ‘the probability of fraud is very low (0.2%) and has been lowered in a preprocessing step by a conventional fraud detecting system down to 0.1%,' while Hassibi (2000) comments that ‘out of some 12 billion transactions made annually, approximately 10 million – or one out of every 1200 transactions – turn out to be fraudulent. Also, 0.04% (4 out of every 10,000) of all monthly active accounts are fraudulent.’ In coping with unbalanced classes, there are two issues to be considered. Firstly, what performance criterion is appropriate? And, secondly, how should the scorecard be constructed, and any parameters estimated, from such unbalanced data? We look at each of these problems. For the first problem, we illustrate the effect that marked lack of balance has on performance criteria, demonstrating how easy it is to be misled. The lack of balance means that simple error counts are inappropriate as performance criteria. Rather, misclassifications of customers from the smaller class must be regarded as more serious than the converse: different costs must be adopted for the two different kinds of misclassification. We examine some of the implications of this. For the second, we examine both classical linear scorecards and more powerful knearest-neighbour nonparametric methods, such as are used in fraud detection. In the case of linear scorecards (and, more generally, for any simple parametric form) improved classification accuracy is achieved by focusing classification performance in particular parts of the data space, with the relevant parts being implied by the relative misclassification costs. We describe a new tool for constructing scorecards which takes this fact into account. In the case of k-nearest-neighbour methods, we draw attention to a phenomenon we believe has not previously been reported, and which has an important effect on choice of k. We illustrate both methods using a large data set of unsecured personal loan data.