Popularity leads to bad habits: Alternatives to “the statistics” routine of significance, “alphabet soup” and dynamite plots

Ruth C. Butler · Annals of Applied Biology · 2021

Combine 3 with 1 in the presentation: Draw a bar chart of the means, with "SEM" error bars. Add stars/letters to each bar OR Table 1a and Figures 1a and 2a show typical examples of results presented using the Routine. These are fabricated, but one or other of these forms of presentation can be found in most issues of almost all biological journals, from the most humble to the most exalted. Other routines (Zuur & Ieno, 2016) are available, but are less widely followed. Example 1a also shows common features: the poor legend wording and no reference to the multiple range tests used (here, Duncan, 1955). This journal (AAB) has some reasonably strong guidelines for statistical presentation: These essentially prohibit use of the Routine. The roots of the Routine are hard to identify, but one of the seeds must surely be Ronald Fisher's book "Statistics for research workers" (Fisher, 1925), where tables of the values of statistics for particular p-values were first published, making statistical testing accessible. In conjunction with this, the last few decades have seen the widespread availability of software that makes the steps in the Routine very easy to carry out, with the consequent explosion in its use. Meanwhile, developments in statistical methodology have proceeded at exponential rates, driven by major developments in theory and computing resources. Whilst still including some significance testing, these newer techniques provide additional and more reliable information. Despite its widespread use, the Routine tends to mean that some useful information in the data is obscured, sometimes leading to unsupported conclusions. There is a very large body of literature, written over several decades, critically addressing the steps, generally suggesting alternative approaches that make more informative use of the data (Vail & Wilkinson, 2020). It would seem that no paper addresses all the steps in the Routine, with many of these papers not published in journals relevant to experimental biologists. This paper provides a partial review of publications relevant to all parts of the Routine. Informative data summary and data analysis are assisted by an understanding of ideas behind the methods used. Therefore, ideas relevant to the steps in the Routine are briefly summarised, along with the suggested alternative practices. The Routine results in a summary of statistical analyses; so understanding the purpose of a statistical analysis is a first step to gaining more information from data. Statistical analysis is about summarising data to find information and about making estimates and predictions (inferences) about what might happen in the future (Fisher, 1925). The need to deal with variation is inherent in these activities. If your dataset includes every possible data point, variation is not a problem because the dataset contains everything: the results of all football matches in a competition are known exactly. Very few research datasets contain all possible data points, so the key problem is to obtain representative summaries, estimates or predictions, and associated measures of uncertainty for these. Significance testing is a part of assessing uncertainty (Fisher, 1925), but has, in the view of many scientists, become synonymous with "the statistics," with the consequence that potentially useful information is not found. An increased awareness of a fuller range of available analysis and data presentation tools can thus lead to an increase in information gained from a dataset. For many biologists, the "p-value" is the entire point of a statistical analysis. Some journals require a capital P, and even that it should be a capital, italicised P, perhaps reflecting this high importance attributed to p-values. AAB uses p, as do the majority of statistical and mathematical publications. There is a dichotomous interpretation of p: A significant "p-value" (usually one less than .05) is interpreted to mean that a difference between treatments is "real." Conversely, if p > .05, then the treatments are "the same." For a given comparison, it is also widely believed there is only one "true value" of p. These ideas are appealing, because they appear to give simple and easily interpretable conclusions. Sadly, the beliefs are mistaken (Goodman, 2008; Sterne, Cox, & Smith, 2001a; Wasserstein, Schirm, & Lazar, 2019). In the context of most significance tests, p simply stands for "probability," and is: The probability of getting a result as extreme or more extreme than the one observed GIVEN THAT The null hypothesis of "no difference" is true AND THAT The assumptions behind the analysis carried out are (sufficiently) true. The last part is often not recognised, but it is very important. p-values can be small because the null hypothesis is not true, or because the result obtained is just one of the rare ones, or because the assumptions about the data required for a valid analysis are not appropriately satisfied. So, for a p-value to be meaningful, the data must sufficiently satisfy the assumptions: where this is not the case, a lack of adherence to assumptions can be the primary cause of a significant result. Statistical significance (or not) says nothing about biological importance (Ziliak & McCloskey, 2008): any "test" result must be interpreted in the context of the trial and relevant biology. Values other than .05 can be used to determine significance. Fisher (1925) promoted choosing a value to reflect the context (e.g., information from previous trials), so sometimes p= .1 might be a good choice, sometimes a value smaller than .05. Fisher suggested p = .05 was convenient, partly because, with the assumption of Normality, it approximately corresponds to a difference being equal to twice its standard error, partly because he thought a 1 in 20 chance was sufficiently small to be interesting (Fisher, 1925). Today, computers can calculate exact p-values in a fraction of a second, so these can be presented, allowing readers to make their own decisions as to the importance of effects (Webster, 2001). A natural interpretation is that the smaller the p-value, the more "evidence" there is that the effect is "real," interpreting 1-p as "evidence against the null hypothesis." Unfortunately, this is not the case. "Evidence against the null hypothesis" could be calculated using Bayes' theorem, but this requires further, usually unavailable, information (Matthews, 2001). In practice, it is reasonable to interpret a large value of p (close to 1) as evidence in support of the null hypothesis (no effect), and very small p values (such as p < .001) as evidence against the hypothesis (Sterne et al., 2001a). Ultimately though, results from many trials and sources are needed to confirm that an effect is "real." Use of asterisks/stars is very common: the practice probably originated as a reference to a footnote referring readers to published tables of critical levels of a statistic. Stars convert a continuous 0 to 1 p-value scale into classes (usually, n.s., *, **, ***), losing both subtlety and information (Wasserstein et al., 2019) whilst simultaneously attributing undue accuracy to p-values near the cut-off points between the classes. It is more informative to present actual p-values, especially for n.s. (not significant), because n.s. can mean anything from a shade larger than .05 (slightly "interesting") all the way up to p = 1 (identical means). p-values are only estimates: analysis assumptions are never exactly satisfied. Thus, very similar p-values, including p = .051 and p = .049, are essentially the "same" (Wasserstein et al., 2019) and should be interpreted as such. In general, p-values should be quoted to only two or three decimal places. One or two decimals are enough for associated statistics (F, t, etc.) if they are also presented. Trial design is fundamental to the value of a study (Finney, 1988), but is not mentioned in the Routine. Frequently, only the number of replicates and treatments used are provided (Haddaway & Verhoeven, 2015), perhaps because it is assumed that the design is always a randomised complete block design. There is a huge range of potential designs, old (Fisher, 1926) to recent (Williams & Piepho, 2019), and sound experimental design underpins effective and efficient studies (Finney, 1988; Kilkenny et al., 2009; Smith & Cullis, 2019). Appropriate analysis methods can vary substantially between different design types. An accurate and complete description of the design is therefore essential, to support the validity of trial results and enable reproducibility (Haddaway & Verhoeven, 2015). ANOVA is the most frequently used method for data summarised using the Routine. Like all statistical methods, ANOVA is underpinned by a set of assumptions that must be satisfied for the results to be valid. ANOVA has five underlying assumptions, with three being the most important (Finney, 1989). The analysis is assessing the differences (meanB−meanA) between treatments rather than other possible relationships (e.g., meanB being a multiple of meanA). The variance is the same for each treatment: Fisher (Fisher & Mackenzie, 1923) developed ANOVA with the primary aim to obtain a more robust estimate of the underlying variation (Finney, 1988). If the variation around each treatment mean is (close to) similar, a pooled (i.e., combined) error can be calculated, based on more data. The pooled error, and any test using it, is more reliable when individually calculated errors are used. Therefore, the pooled measure of error should be used in the presentation of analysis results (Welham, Gezan, Clark, & Mead, 2015, chap. 5). Any data point is independent of another data point. Lack of independence is often caused by taking multiple measurements of the same thing, either at different times ("Repeated measures") or the same time ("Pseudo replication"). The analysis needs to be adjusted to allow for this, otherwise test results will be unreliable, frequently with p-values being too small. The remaining assumptions are less important. The data around each mean is Normally distributed. Normality underlies the validity of the testing of the F or t statistics, although not always the validity of the statistics themselves (Finney, 1989). ANOVA is "robust" to departures from Normality (Lumley, Diehr, Emerson, & Chen, 2002), whereas it is not strongly robust to variance heterogeneity (Welham et al., 2015). Bias is when the result is skewed because of experimental procedures, such as the use of a systematic layout, where treatments have not been randomly allocated (such layouts can sometimes be justified). Assumptions need to be checked, and only need to be approximately satisfied. Assumption checking however is frequently confined to assessing Normality, or not checking assumptions at all (Warton & Hui, 2011). Where assumptions are not sufficiently satisfied, many scientists address the problem by using modern statistical methods. However, the two most common ways to address violations of assumptions are to transform the data and then use the same analysis method, or to use a nonparametric method. Each has drawbacks. ANOVA and standard regression are parametric methods because the primary output from the analyses is estimates of parameters (predicted means, regression parameters, etc.). Parametric techniques involve fairly strong assumptions, whereas nonparametric analyses require much less stringent assumptions to be satisfied. The most commonly used nonparametric methods (e.g., Mann–Whitney U) obtain ranks of the data, and then carry out tests on those ranks rather than on the actual data. Thus, the primary output from these methods is not estimates like treatment means, but a test statistic, assessing the similarity of the mean of the ranks for treatments. p-values for these statistics will usually be larger than those for the same comparison from a parametric analysis, because more assumptions are made for a parametric analysis. Thus, in general, the parametric test has more "power." The assessment of differences between mean ranks does not automatically apply to differences between treatment means: care needs to be taken when interpreting the results of a nonparametric analysis. Historically, data transformation prior to analysis was used to enable the assumptions behind standard ANOVA or regression to be sufficiently satisfied, thus enabling these methods to be used. The primary reason to transform data was a lack of variance homogeneity and to allow analysis of non-Normal data (Faraway, 2002; Welham et al., 2015). In an age with only rudimentary computing power, this was the best that could be reasonably done. Nowadays, methods that more directly model the characteristics of the data are available and should be used in preference to data transformation (Lane, 2002). A key advantage is that the results are then generally more interpretable from a biological point of view. A transformation modifies several things (Lane, 2002) that need to be considered when interpreting analysis results. These modifications are illustrated with a log transform. If the data for each treatment are log-Normally distributed, then the distribution of the logged data is Normally distributed (Evans, Hastings, & Peacock, 1993). For log-Normally distributed data, the variance increases with the size of the treatment means. The variance of the log(data) is often constant across treatments. ANOVA of the log(data) gives the means of the log of the data, which when back-transformed are geometric means. For n data points, the geometric mean is the nth root (i.e., to the power of 1/n) of the product of all of the data points. The geometric mean can be quite different from the usual (arithmetic) mean: The mean of 5, 9, 12, 25, 27 is 15.6 and the geometric mean is (5 × 9 × 12 × 25 × 27)1/5 = 12.95. Means and differences between means obtained on the log scale can be back-transformed, as can confidence limits or a least significant difference (LSD), which becomes a ratio. (LSD: the smallest difference between two means such that the means are significantly different at the required significance level.) However, standard errors cannot meaningfully be back-transformed. ANOVA of the log(data) is comparing the differences between the means of the log-data. Back-transforming this difference gives the ratio between the geometric means. This does not directly say anything about differences between the raw means: Untransformed data: MeanA−MeanB Logged data: log(GeoMeanA) − log(GeoMeanB) which, when back transformed, is: GeoMeanA/GeoMeanB. Historically, the square root transform was used for count data, and arcsine for proportions out of a fixed total. These both change the However, for the log there is no way to the difference between the means or and for the using ANOVA to these gives results. There are where transformation makes such as a square root transform for or root for to a Welham et an case, a transformation will give interpretable of the as as enabling it to satisfy the assumptions of the analysis so that the are Where the data do not sufficiently to the assumptions, data transformation and nonparametric methods can generally be because methods are available that allow the characteristics of the data as to be methods are in some papers in many journals and are in many (Welham et al., 2015). & provide methods for the analysis of both and proportions with interpretable results. the use of for is very but to for Normally distributed data, which is just one of In this always to a For a analysis, means are to the raw means, and the between and the means is by the The is an part of a For a log and a simple model with one and associated to = Other useful analysis methods are for with fixed and effects (Welham et al., 2015), or of a model & 1993). If a dataset is using more than one method, a range of p-values will be obtained for comparing the same and can The p-value for a comparison is by several including the of comparison between treatments between means, ratio of means, number of and analysis of the data in Table 2a test results can change between different methods. The data are for five replicates of two A and Table has the summary The treatments were using different but valid with p-values as The first test is a standard using the pooled The is the same but with the p-value calculated using a making the test partly In this all values of the data are many each time five data values to each The p-value then extreme the result with the actual data was to the results for all of the data. The test is the Mann–Whitney a nonparametric of a This test the means of the ranks of the data & 1988). The test the treatments using a The model includes a log the to the means, which to the ratio of the means being rather than the this is a change in the hypothesis being The p-values vary from for the standard to for the The two nonparametric tests have similar, but not p-values, which are larger than those for the parametric this is because of the assumptions underlying the parametric The interpretation of the test result with the because the treatments are being in different It would be to that the treatment means vary significantly because the result of the Mann–Whitney test was However, can say that the values for treatment A were than those for the ratio of the treatment means, as in the does not much about the difference − between the means. as test results can vary between different methods, so can errors and significance results from one analysis method are not generally to those from another method. analysis methods need to be and the interpretation and presentation of results will be by that method. Thus, it is usually not to results from different analyses methods, including simple summary statistics from the results of a analysis. A trial or other of study has that determine the study is carried out & Mead, & The study lead to to be which lead to the of to be treatments and treatment & The statistical analysis including tests at least in be even the data are Thus, the need to each treatment with each other treatment should be very testing (or is a of the Routine, leading to However, the results are often not easily interpretable or a with treatments different of a The would be is the between and the of and is to be the possible of which are significantly different from each Where there is a large number of treatments (e.g., several it is often useful to make an test (e.g., in and then present the means in of along with a measure of This can be the means in This method can in the treatments. The aim for is often to find the few so testing every of treatments is often not Some are of more than (Finney, is the comparison of the with the and of as much as the one with the of comparison are generally presented using Figures 1a and sometimes important in the data. However, they convert a continuous 0 to 1 p-value scale into two classes and not significant), losing even more information than the difference" "the also the key information between estimates of In if the were not the for would be is than and which are reasonably The nothing they that is "the as but to which is "the as It is more to an error such as the as in For the make interpretation rather can and be different from whilst and are "the The effect seen can be a result of or confidence are more The with which all the treatment can be an analysis substantially between statistical some the only way to make the required be to do that use the standard error of a difference or an from this way will give p-values to the same an it is to make treatment in such and that the tests all should generally be as with a view to to be in future analysis it is most useful to on the of treatment in the data, especially those of biological with actual p-values used as just the p-value or was is almost effects can be as interesting as significant so can be et using the A major for presentation of results is to There are many ways of the as to which is most useful between show the Routine uses either individually calculated "SEM" or The is a & & & Wilkinson, and is the of the from the so the variation in the data For Normally distributed data, on of the data will from the mean & & In the "SEM" is an and is a rudimentary analysis (Vail & Wilkinson, 2020). It is a measure the mean of a to the mean of the underlying the validity of an "SEM" is on to analysis assumptions such as Normality to have any The "SEM" calculated individually for a treatment gives a of the variation in the raw data, because it is always smaller than the an estimate of variation or of a the "SEM" is as is it is calculated from a small number of values 3 or The is in the large associated the t = and for 3 and ANOVA addresses this by information over the the pooled standard error has larger associated of and smaller so is to individually calculated the is the "SEM" so widely The would appear to what always been that the "SEM" is smaller than the so it and a that the "SEM" can be easily used to two means are and on scientists cannot say means need to be for "SEM" to allow of making rather In to the and there are many other ways of There are two & variation in the raw data, and of estimates as from an analysis. Appropriate ways to show vary between the two so the of tables and also The of error are widely 2009; & & & Table 3 the types. of error are in a or the error needs to be in the because interpretation of the information substantially between the error statistics are in the same way as error to but they are of error give quite different and so must make what error & In the use of to a measure of error is and can be & it a range that the can vary just so it should generally be In errors should be for use = = = It is frequently useful to the with an error & & because they give a measure of of the error, and can allow it to be into an alternative error (e.g., convert into confidence In of data are summarised 1 of & 2011). The mean and are exactly the same for each of the = = Thus, the bar of the mean with are all The bar with "SEM" only because the errors are calculated individually for each so are where n = the is driven by not by the variation in the data or are widely by the statistical & & Wilkinson, because they many of the in the data. The more presentation of the data, as illustrated with with the from and that the both in the distribution of the data and the number of For both values of 1 is from a which is is from a so is 3 is an set where the data are in two For a summary of the raw data, is the most and show of the information that can be seen in Where n is too large for presentation as in or such as & be However, the standard does not the seen for 3 in the some show that the "SEM" does not reflect the variation in the raw data, it is more informative to present the in for non-Normal data, present the and of the data or the range & & For or & & Wilkinson, et al., 2019) or use less than the standard bar chart with error whilst more information information to values are parametric analyses carried out with sound statistical software calculate estimates and errors associated with as a part of the analysis, because these is the primary reason to carry out the analysis. Thus, these estimates and errors should be presented, including when the estimates are to the raw means. In it can be to show statistics results from analysis. However, it is less to the two The of errors presented should be on the and of the the of the data and what are to In a confidence limits on each mean be most If have carried out ANOVA on the raw data, the means and an be especially where the is the same for all treatments. the for the examples in Table 1a and Figures 1a and 2a are in Table and Figures and The bar chart is the very many other different ways to data as a and the large presentation will show et al., 2019). The of bar to be partly because that is what partly a lack of with other and partly a that a bar a in bar are both the and to Whilst bar can be a useful way to present data, other will often be more three to bar are presented a a sometimes a and a which is also a or The first two are generally most useful for the results from an analysis, and the for data these only the can easily be made with However, can be obtained to the other Statistical software etc.) and software etc.) enable these other to be presentation and both in choosing what the should and the it, such as and It is important to about what are to can be do not just the first out several different and different software the of available and some make it very hard to the a statistical as part of your Where this is not with an a research with rather than treatments that can this and which treatment are of Use a sound trial the trial layout, etc.) your data using summary statistics and out a sound statistical analysis, using an analysis method that as much as is the characteristics of the data. estimates and associated measures of from that analysis the treatments as by their the and the research the trial design and statistical analysis methods. If present raw data results from the estimates and associated measures of the informative parts of your data and analysis by including tables and types. that all parts of the p-values as evidence or interpret the the interesting of treatment in the data. Use of alternative to the standard Routine often more and for This is at least partly because software makes using the Routine However, the to more informative practice are not some to more about statistical methods, and to use a range of The more and information for the same in data. There is a that is relevant to the in this This includes several of additional to those in the were used to the ideas presented in this

Read the paper · More papers on PaperTik