Scorecard Development Process, Stage 4
Naeem Siddiqi · 2016
This chapter deals exclusively with model development using grouped attributes and logistic regression. Simple statistics such as distributions of values, mean or median, proportion missing, and range of values for each characteristic can offer great insight into the business, and reviewing them is a good exercise for checking data integrity. Some data mining software, such as SAS Enterprise Miner, contain algorithms to impute missing data. Such algorithms include tree-based imputation and replacement with mean or median values. Multicollinearity (MC) is not a significant concern when developing models for predictive purposes with large data sets. The effects of MC in reducing the statistical power of a model can be overcome by using a large enough sample such that the separate effects of each input can still be reliably estimate. The R-squared technique uses a stepwise selection method that rejects characteristics that do not meet incremental R-square increase cutoffs. One of the major causes of biases in banking data for application scoring is lending policies and adjudication based overrides.