The impact of experimentation guides in software engineering: a metanalysis
Italo Macêdo do Amaral Costa · 2015
Empirical studies are needed to develop or improve processes, methods and tools for software development and maintenance. There are several scientific articles and books dedicated to empirical software engineering methodology. From now on, we’re calling them experimentation guides. The literature shows a lack of consensus for the process of experimenting. Thus, there’s a need for evaluation on the experimentation process itself. Also, the current state is particularly worrying. There is body of knowledge with a low statistical power, poor validation and a few replications. In turn, this might lead to deficiencies in the accumulation of knowledge and the presentation of advice to industry. In a preliminary non-systematic literature review, We found several secondary studies that evaluate the software engineering domain in the following aspects: Statistical power and validation. Thus, we decided to measure the impact of experimentation guides in software engineering. In order to measure the impact, 4 hypothesized conclusions were taylored and grouped in two stratum: the validation stratum, in terms of the number of threats (H1), mitigative actions(H2) and the mitigation rate of each study(H3); and statistical power, in terms of the statistical power itself (H4). For comparison purposes, controlled experiments were grouped considering the experimentation guide referenced or not by each experiment. 133 studies ranging from 2002 to 2014 were included in the metanalysis. 69 studies mentioned a experimentation guide while 64 studies did not mention a experimentaiton guide. Overall, 722 threats, 402 mitigative actions and 266 statistical tests were retrieved. For the validation stratum, all the hypothesized conclusions showed a difference between the two groups that is statistically significant. Studies that reference a experimentation guide tend to identify and treat more threats, also they showed a higher rate of mitigation per study. For the statistical power stratum, overall differences between the two groups were nonexistent. However, when comparing studies that did not mention a experimentation guide with studies that referenced Basili et al. the difference was still statistically significant. Studies that referenced Basili et al. achieved a statistical power higher than the experiments among the control group.