The New Statistics with R: an Introduction for Biologists. — By Andy Hector.

Krzysztof Bartoszek · Systematic Biology · 2015

The R language (R Core Team 2013) is growing more and more popular with researchers, and naturally more and more books about it are appearing. Its open-source nature allows for quick development even of very complex models and their application to data, especially in the natural sciences (in the evolutionary biology community, see Nunn 2011; Paradis 2012). The New Statistics with R by Prof. Andy Hector is a nice addition to the R community; and this review provides an opportunity to discuss statistics teaching in biology, and to present my own (very) subjective list of useful R books. This results from my research and teaching of statistics and mathematics, and their applications to life science data, nearly always using R. Therefore, this book is of great interest to me, and it is important to evaluate its value as a teaching tool for R for biologists. R was created as an open-source freeware implementation of the S computer language—a rapid development programming language for statistics developed at Bell Laboratories in the mid-1970s. (The commercial version of S is called S-PLUS.) As this is its intended application, R is very widely used in statistics classes, both for introductory courses and specialized advanced courses or workshops. Naturally, this means that many textbooks connected with R are now available on the market. Probably the best known English language is one that of Dalgaard (2008). However, it seems that standard undergraduate statistics course books (or at least those that I have worked with) use the Matlab, SAS or SPSS programs, instead. This is a consequence of their first editions appearing way before the start of this century, which is when R started gaining popularity (first released in 1993). If R's popularity continues to grow, then we may expect to see it “taking over” undergraduate textbooks—already, free online and more advanced texts (like Gelman and Hill 2006) tend to prefer it. The existence of a flexible programming language tailored for statistics and widely available computational power may, of course, lead to the firm establishment of a new teaching paradigm—statistics is taught not through formal derivations but rather through the students' experimentation with data and models. This teaching approach already exists, of course, as a complement to the traditional approach. Andy Hector's The New Statistics with R, on the other hand, nearly completely replaces the formal approach with informal (but of course carefully prepared and presented by the author) experimentation with data by the student; in the author's own words: “the approach is to learn by doing through the analysis of real data sets.” The book is a course in statistics for an applied scientist, and is essentially a compressed version of Dalgaard's (2008) work. The author begins with ANOVA (Chapter 2), t–tests (Chapter 3), and linear regression (Chapter 4). He then moves on to factor analysis (Chapter 6), ANCOVA (Chapter 7), maximum likelihood (Chapter 8), generalized linear models (Chapters 8 and 9), mixed-effects (Chapter 10), and generalized mixed-effects models (Chapter 11). A key chapter is Chapter 5 (Comparisons using estimates and intervals). In this chapter, the author discusses his inference philosophy—to focus on effect sizes and looking at values of estimates rather than on P-values (these are very nicely explained in Box 2.7). A welcome part of the book is that the author explains information criteria in a very accessible way. Naturally, the book does not cover everything there is concerning statistics and R. For example, R for Bayesian statistics is covered by Albert (2007), Cowles (2013), and Marin and Robert (2013); multivariate statistics is covered by Everitt and Hothorn (2011) and Bilder and Loughin (2014) discuss categorical data, devoting a lot of space to model selection and evaluation (although I personally prefer the book by Gatnar and Walesiak 2011 in this respect, as they also include cluster analysis and classification trees). Prof. Hector's book should be very much welcome by students, statistics teachers and practitioners. The individual chapters are ready-made computer laboratory sessions. Working through them will certainly teach the basics of data analysis. Hector's informal approach to statistics is, from my experience, the one desired by applied researchers. In the book, statistical inference procedures are often treated as black boxes, informally discussed in separate parts of the text—for example: “I have minimized the number of equations—they are in numerous statistics textbook if you want them ….” The researcher's main responsibility is to correctly identify the response and predictor variables, and then to choose the correct analysis method. This is in line with the author's stated philosophy: “for non–statisticians the analysis should not become an end in its own right, only a method to help advance our science.” However, in practice it turns out that such an informal approach can sometimes have short legs. The choice (and justification!) of the correct method requires the skill of probabilistic modeling. This is not an issue of mathematical elegance, but of correctly handling data and interpreting the information contained in them. From my own experience, an analysis is unfortunately hardly ever as straightforward as comparing groups with a t-test. Actually, in not a single one of the studies that I have been involved in were standard analysis methods satisfactory. In no case could I just read in the data and run a single statistical R command on the columns of the data frame. Indeed, whole research areas can be built around such shortcomings of standard approaches. For example, in phylogenetic comparative methods one of the fundamental assumptions is immediately violated—the data are not independent. Therefore, inference methods have to be adapted and appropriate stochastic models developed to deal with this situation. It is, furthermore, difficult to explain inference methods without the underlying mathematics. Even in The New Statistics with R this is evident. For example, to estimate the grand mean one can use lm(y ∼ 1), or to compare two samples we use the t-test for the difference of the means. These are presented in the book with an intuitive justification, but the student could be in a lot of trouble if asked “why not lm(y ∼ 2), or why not use the quotient of the means?” Fortunately, everything in the book is well illustrated with R code, so hopefully the reader will be able to follow up on many of the “whys.” The sections that I found most difficult to read were those concerning Generalized Linear Models (GLMs), mixed models and Generalized Linear Mixed Models (GLMMs). These complex models are very accessibly discussed, but the lack of any underlying equations requires constant access to the background literature. As already mentioned, the main point Hector makes in his book is that instead of reporting P-values, the researcher should concentrate on discussing confidence intervals (CIs) and effect sizes — “I encourage you to attempt the P–free challenge.” A second point is: “I avoid the use of corrections for multiple comparisons (and discourage their use in many cases).” These two statements illustrate very well, in my opinion, the problem that arises when teaching statistics “in a practical way.” In a way, it seems that Hector is giving up (by referring interested readers to literature) on explaining correct handling of these more difficult statistical procedures; and, rather, he says stick to this simple set of tools—you have a lesser chance of getting it wrong with them. The author gives the classical example of “marginally significant” and “non-significant,” arguing that it is better to use CIs as they will not have such a problem attached. However, when constructing a CI one has to choose a confidence level (the same as the a priori choice of a significance level), and this will determine the width of the interval. And in the aforementioned case we will run into the same problem—a slight alteration of the confidence level will change whether the value under the null hypothesis is covered or not covered by the interval. This is well illustrated by Fig. 5.4 C in the book (which has a very wide CI) with the author's comment on it: “while this result is non–significant by the convention of P < 0.05, there is a possibility of a large effect ….” But the main conclusion from such a CI is in the following paragraphs, that one should work on the data collection—as Prof. Aaron King very nicely summarized: “To put it another way, if you are asking a question that hinges on the value of a parameter, and the CI for that parameter is so wide that your question goes unanswered, this is because the data do not contain an answer to your question” (quoted from the R-sig-phylo Digest, Volume 86, Issue 23). Rules of thumb, such as report CIs instead of P-values (or vice versa) or avoid multiple testing procedures (or the opposite), can blur what these objects really are. A CI is better than a P-value, in the sense that it contains more information—both the significance (at the prescribed level) and the magnitude of the effect. But then drawing a conclusion from a CI that something is important (i.e., significant) is no different from using P-values. The mathematics in both cases is consistent, and so both have to result in the same call. Therefore, I think the focus should be on teaching what these objects are. Subsequently, the choice of which to use should depend on which one will better convey the science in the study. Otherwise, without knowledge of the underlying theory one hears (as I have): “I will rather stick to classical t–tests instead of Bayesian statistics as I don't want to make distributional assumptions”—but t-tests assume a normal distribution (robustness is a separate issue)! I agree with Prof. Hector that there are many cases where it is advisable not to use multiple testing corrections, especially as P-values can be dependent—for example, the Bonferroni correction is too conservative. But, to make the call to use or not use them, one has to be able to justify the decision; for example, that there are strong dependencies between the P-values, or that despite the low significance (wide CI) there is a mechanism justifying the effect. The risk with “rules of thumb” statistics is that many people will stop at these, and not make the extra effort of investigating and strengthening the evidence for weak signals. In consequence, they will inflate the number of false-positive findings, a common issue in, for example, biomedical sciences, so that subsequent translational or replication studies do not confirm the results. Prof. Hector rightly emphasizes that the only way to eliminate such weak-signal false positives is to repeat the experiments. However, as Schlattmann and Dirnagl (2010) point out: “Only by understanding the basics of a statistical test will the researcher be able to use it properly and interpret the results given by computer programs.” In their paper they reiterate that one needs to do an ANOVA (very nicely covered in the second chapter of Hector's book) instead of all possible pairwise t-tests, as “failure to adjust for multiple comparisons is highly prevalent in many fields.” I am not implying in any way that I think that Hector is downplaying probability and statistics. Rather, the teaching approach in his book is only one particular way of doing it. Of course, Hector's lengthy research experience gives him the ability to pick out the statistical inference procedures that are useful for an applied researcher. And, by the way, he makes a very important point in his book: “a linear regression model that explains much of the variation of the data” can be established despite many potential issues about the assumptions. This is similar to what I often suggest—experiment with lm() and, unless you are extremely unlucky, it will still catch the most important relationships in the data, and indicate further steps for study. However, often, to be able to decide upon the contents of one's toolbox one needs to know all of the details about the tools. Not only what is useful, but why and how it is useful, and in what ways the tools can be put together to create new ones or be used in atypical situations. Therefore, one needs a mix of both approaches—knowledge of what to apply, and the mathematics behind it. This, of course, raises the question of whether there is any book concerning statistics and R that encompasses all of the necessary components—programming, statistics, graphics, and mathematical background. Dalgaard (2008) essentially does this, and hence his textbook is a standard one, at the moment. However, he does not cover model selection, or more advanced programming techniques. The former is also not covered (actually no programming is) by The New Statistics in R—in it, the focus is on single commands not complex sets of commands. Crawley (2009) wrote the traditional “everything about R” book—but the nearly thousand pages may seem daunting. I found that Biecek's (2011) book best balances between all four topics. Of course, this opinion can be due to an educational bias—it is written in the style used in my own university studies. Apart from standard inference procedures, Biecek (2011) discusses in very much detail (surprisingly, maybe even more than Crawley 2009 and Dalgaard 2008) R's graphics possibilities. Again, probably due to my educational bias, Ga̧golewski's (2014)'s book seems most useful for me concerning advanced R programming (environments, code profiling, S3 and S4 systems) and scientific computations (optimization, interpolation, solving systems of equations). Unfortunately, Biecek's and Ga̧golewski's books are currently available only in Polish; and therefore they will not be useful for the majority of readers. For English speakers, Dalgaard (2008) book and The New Statistics with R are good places to begin. Dalgaard (2008) covers the most important material, and Hector brilliantly illustrates it with data analysis and intuitive explanations. A more programmatically oriented reader might find it useful to have at hand some position on advanced R techniques (e.g. Burns 2012; Wickham 2014). The reviewed book is written by a very experienced practitioner. Despite the risk of misuse of the author's approach (as discussed above), the book's strength is that it takes an applied scientist through the necessary basic statistics, and shows step by step how to work with real data. The New Statistics with R is, furthermore, a great textbook for computer exercise sessions in any introductory statistical class (especially for the life sciences). With its help, one should be able to design a very attractive course for both applied and more theoretical students. I would like to thank Anna Stokowska for many discussions and insights concerning the use of R with biological and medical data, and for substantial comments on a preliminary version of this text.

Read the paper · More papers on PaperTik