Some comments on the update to BJP guidance on experimental design and analysis
Markus Neuhäuser, Graeme D. Ruxton · British Journal of Pharmacology · 2018
We support the development of guidance for experimental design, analysis and reporting (Curtis et al., 2015, 2018) to improve rigour and reproducibility. However, from a statistical point of view, some issues of the updated guidance (Curtis et al., 2018) may not improve on previous guidance and may even be a step in the wrong direction. Fortunately, we feel these mis-steps could be easily corrected. According to the 2015 guidance, 'when small groups (n < 20) are used, they should be of equal size unless a valid scientific justification for unequal group sizes is provided' (Curtis et al., 2015). The updated guidance says 'Studies should be designed to generate groups of equal size … (with credible justification if not possible)' (Curtis et al., 2018). Thus, the updated guidance implies that equal-sized groups are the 'gold standard', but there are several situations in which the careful researcher might be justified in planning to have unequal sample sizes. When k different treatment groups are compared with a control and the well-known Dunnett procedure is applied, the square-root sampling allocation rule (n for each treatment group, n0 = n for the control group) gives a more powerful test, or the same power with a lower total sample size, than equal group sizes (Liu 1997). Consequently, when there are N subjects in total, maximum power is obtained with unequal group sizes. The situation where several groups are compared with a control occurs in the 'majority of papers published in BJP' (Curtis et al., 2018). Thus, it is counterproductive to encourage equal group sizes so strongly. A further reason for unequal group sizes is unequal variances. To increase power, a larger proportion of the total sample size should be allocated to a group with a larger variance. In this situation, again, maximum power is obtained with unequal group sizes when there are N subjects in total. Moreover, if one treatment involves more potential for suffering (or higher expense), an unbalanced design may offer attraction. For instance, some placebo-controlled trials uses a 1:2 or another unbalanced randomization in order to treat fewer patients with placebo. Thus, we feel the guidance could be improved upon by encouraging researchers to consider planning unbalanced designs more readily. The 2015 guidance stated 'a priori sample size calculation … should include alpha, power and effect size' (Curtis et al., 2015). Now, it is recommended to increase the sample size determined by power analysis by 50%, without suggesting any desired value for the power (Curtis et al., 2018). We agree that when the power is not very high, there is a high risk of false negative findings (type II errors). But a general principle to 'add 50% to the calculated minimum group sizes' (Curtis et al., 2018) is unusual and not reasonable. To reduce the risk of false negative findings, one can increase the desired power. A percentage-wise increase of the sample size gives no control of the actual power because the increase in power depends on the chosen statistical method and the given assumptions. Thus, statistically sound advice would be to demand a minimum power of e.g. 80% and encourage researchers to seek higher values such as 90% or 95%. The blanket suggestion to perform a power analysis and then increase the calculated sample size by 50% is unjustified. Rather researchers should be encouraged to consider costs and benefits of different desired levels of power and then calculate the sample size required to give that power. We agree that significance in classical ANOVA can be caused by inhomogeneity in variances. But the recommendation not to carry out post hoc tests in case of a significant variance inhomogeneity is not satisfying. A better strategy would be performing an ANOVA designed for possible variance inhomogeneity. Several methods have been proposed for this (e.g. Brunner et al., 1997). The guidance sets '5 as the minimum "n" required for datasets subjected to statistical analysis … any data set containing groups of n < 5 must not be subjected to statistical analysis' (Curtis et al., 2018). Clearly, asymptotic or approximate tests are not acceptable for very small samples sizes, but the minimum 5 is completely arbitrary. A reasonable minimum might depend on the type of the endpoint, the number of groups and the chosen statistical method. Inference based on small samples might be delicate, but a statistical analysis with group sizes of n < 5 should be presentable if there is a scientifically sound explanation why this analysis seems to be reliable despite the small sample. Fisher's exact test, for instance, can attain a two-sided P-value as small as 0.029 with n = 4 per group. Moreover, special statistical tests have been developed able to attain small P-values even with very small group sizes (Beasley et al., 2004). These tests are useful when an observed difference is huge. The authors declare no conflicts of interest.