The Effect of Sample Size on Different Failure Prediction Methods
Barbro Back, Teija Laitinen, Jukkapekka Hekanaho, Kaisa Sere · 1998
Neural networks and machine learning methods have proved in many ways and in a number of publications to be real challengers to statistical methods - especially to logit and discriminant analysis - in predicting failures. However, most of the studies have used a rather small data set, very often close to only one hundred observations. Therefore, it has been difficult to say whether there are any significant differences between the methods tested. In this study, we compare neural networks, a machine learning method, discriminant analysis and logit analysis using a large data set consisting of 570 companies. We investigate the effects of the prediction capabilities of the methods when using different sample sizes for estimation and testing, i.e. 400-170, 200-90 and 100-50. Our study shows that neural networks and the machine leanring method perform better than discriminant analysis and logit analysis when the sample size is 400 while there is no best performer when the sample size is decreased to 200 and 100.