On the role of statistical significance in exploratory data analysis
Inderpal Bhandari, S. Biyani · 1994
Recently, an approach to knowledge discovery, called Attribute Focusing, has been used by software development teams to discover such knowledge from categorical defect data as allows them to improve their process of software development in real time. This feedback is provided by computing the difference of the observed proportion within a selected category from an expected proportion for that category, and then, by studying those differences to identify their causes in the context of the process and product. In this paper, we consider the possibility that some differences may simply have occurred by chance, i.e., as a consequence of some random effect in the process generating the data. We develop an approach based on statistical significance to identify such differences. Preliminary, empirical results are presented which indicate that knowledge of statistical significance should be used carefully when selecting differences to be studied to identify causes. Conventional wisdom would suggest that all differences that lie beyond some small level of statistical significance be eliminated from consideration. Our results show that such elimination is not a good idea. They also show that information on statistical significance can be useful in the process of identifying a cause.