Statistical significance in empirical software engineering research: A maturity model
Richard Torkar, Francisco Gomes de Oliveira Neto, ROBERT H. FELDT, Lucas Gren, Carlo A. Furia · arXiv (Cornell University) · 2017
Software engineering research is maturing and papers increasingly support their arguments with empirical data from a multitude of sources, using statistical tests to judge if and to what degree empirical evidence supports their hypotheses. The objective of this paper was to develop a statistical maturity model. First, we manually reviewed papers from five well-known and top-ranked journals producing a review protocol along with the view of current (2015) state of art concerning statistical maturity, practical significance and reproducibility of empirical software engineering research. Our protocol was then used as ground truth, i.e., training set, for a semi-automatic classification of studies for the years 2001--2015 using close to 4,900 papers. We used the extracted data to develop a statistical maturity model. The statistical maturity of empirical software engineering research has an upward trend in certain areas (e.g., use of nonparametric statistics, but also more generally in the usage of quantitative analysis). However, we also see how our research area currently often fails to connect statistical analysis to practical significance. We believe that the statistical maturity model can be used by researchers and practitioners to build a coherent statistical analysis and guide them in the choice of statistical approaches of its steps.