How to Be a Better Scientist: Statistical Training

Christopher T. Filstrup · Limnology and Oceanography Bulletin · 2020

This summer, I had the pleasure to discuss science with educators participating in the Trimming our Sails Workshop organized through Minnesota and Wisconsin Sea Grant and the Center for Great Lakes Literacy (https://www.cgll.org). For the workshop, I was asked what I thought were the most critical skills for students interested in a science career. My response was science communication and statistics. If you have previously read this column (or the L&O Bulletin for that matter), then you are likely already aware of my thoughts on the former. The latter has been on my mind these days as I have a new graduate student that just started in my lab who has limited statistical training (certainly not surprising for recent undergraduates). I wanted to help her avoid some of the challenges that I faced while improving my statistical toolbox. My statistical journey went something like this. I had limited exposure to statistical methods as an undergraduate and only picked up the basics (ANOVAs anyone?). In graduate school, I took the required introductory statistics course, which was taught in another department, and learned about study design, more on ANOVAs (which now made sense), and regression techniques. I felt that I really learned stats as I began analyzing data for my dissertation and for publications during my postdoc. During this journey, I was taught in one program, switched to another program for advanced statistical methods, and then switched to another program when I moved for my postdoc due to institutional licenses. And this does not include the different programs that were needed to actually visualize relationships. Then, I learned R (R Core Team 2019) during my postdoc. I must admit that I resisted learning programming in R at first. The learning curve seemed too steep, and who has the time? I was “encouraged” to use R as I began to work on a large multi-institutional research project where it seemed that everyone had a much better statistical background than I did. I had trouble following along with some of the research approaches and analyses because I could not understand the programming language. Luckily, I had very knowledgeable (and patient) collaborators to help me on my journey. I spent my time buried in Crawley's R Book (Crawley 2007; now in 2nd edition) and searching the online forums to fix coding errors. All of this seems like a distant memory, and I am now happy to spend a half-day or full-day writing code. Fortunately, the new generation of scientists will not have to experience such hardships as accessibility to statistical training resources has increased greatly. As I typed that, I felt a cold shiver as I know that someone out there is reminiscing about punch cards and the fact that my generation actually had computer programs (and computers that fit on a desk) to run our stats. I reached out to a couple of colleagues for recommendations on statistical training modules, largely focused on the use of R, which follow below. Interactive lessons in the swirl software package (https://swirlstats.com), recommended by Cayelan Carey, Virginia Tech. SWIRL is an R package that turns the RStudio console into a learning environment in which you can complete different predeveloped lessons for learning R features. Programming with R materials from Software Carpentry (https://swcarpentry.github.io/r-novice-inflammation), recommended by Joe Stachelek, University of Wisconsin-Madison. Unlike some other resources, the materials are not designed to be the most concise overview of R for people who are already programming masters. Instead, they are designed to teach you the basic concepts that all programming depends on and learning R is simply a happy side effect. In addition, the materials are actively maintained by the great Software Carpentry community following evidence-based best-practices of teaching. Ecology curriculum lessons from Data Carpentry (https://datacarpentry.org/lessons/#ecology-workshop), recommended by Cayelan Carey, Virginia Tech. These also have Python equivalents and are specifically geared toward data analysis, visualization, and workflows. Modules from Macrosystems EDDIE (https://serc.carleton.edu/eddie/macrosystems/modules), recommended by Cayelan Carey, Virginia Tech. All of our teaching modules are geared towards beginners (and instructors) who have not worked in R before; students work through predeveloped R code that they modify to model lake ecosystems and study how lakes respond to climate and land use change. Each module is plug and play and designed for a 1–3-h lab period at the undergrad level. If you are interested in incorporating Macrosystems EDDIE into your course content or to simply learn more about the modules, Farrell and Carey (2018) describe their experiences and lessons learned while developing an undergraduate curriculum in environmental data-driven approaches. Happy coding, everyone! If you have any thoughts or critiques of this column, or suggestions for additional resources on this topic or for topics to cover in future issues, please feel free to contact me. You can email me at [email protected] or tweet me @ctfilstrup. Please be sure to tag the L&O Bulletin (#ASLO_Bulletin) in your tweets.

Read the paper · More papers on PaperTik