Appropriate use of p‐values

Michael Jones, Kirsteen N. Browning, Maura Corsetti, Dániel Keszthelyi, Andrea Shin, Rajan Singh, Pierfrancesco Visaggi, Frank Zerbib · Neurogastroenterology & Motility · 2023

The use of statistics, particularly inferential statistics, is ubiquitous in most areas of medical research. Because we usually only have data from samples, and sometimes quite small samples, the ability to differentiate treatment benefits or risk factor associations that might be reproduced in the broader population of patients from those that may not is crucial. The article by Bangdiwala raises some important issues in how the practice of statistical data analysis frequently differs from how it should be used. Statistical hypothesis tests are an important tool in medical research and when used as intended and in a targeted fashion, the benefits are enormous but misused or overused can give a false confidence in findings that do not actually support the researcher's scientific hypotheses. As Bangidwala points out, the statistical hypothesis testing tools that we use today were developed a long time ago and in a very different context, and mostly nonmedical contexts. The purpose of statistical hypothesis testing then was to test very specific hypotheses, and usually just one. In some areas of medical research, particularly basic science, that can still be the case. In many areas, however, a single manuscript will report multiple effect sizes that represent differences between groups or associations, each with an associated p-value. The intention is good, to give the reader some indication of whether the null hypothesis should be rejected or not. However the discrepancy between principle and practice of statistical hypothesis testing means that they cannot always be taken at face value. The editors of this journal feel that abandoning the use of statistical hypothesis tests would also be harmful, as they do serve a useful purpose. We support the suggestion of Bangdiwala that statistical hypothesis tests be used in a targeted fashion rather than indiscriminately and, where possible, are replaced by other measures, such as interpreting the width of confidence intervals and whether the confidence interval includes the null effect. We also argue that authors should declare whether their use of statistical hypothesis tests is literally to test an a priori hypotheses or is used in a more exploratory way. There is absolutely nothing wrong with exploratory data analysis, but it should be declared rather than dressed up as true hypothesis testing. Michael Jones contributed to the manuscript writing, co-conceived of idea, and approved final manuscript. Kirsteen Browning, Daniel Keszthelyl, Andrea Shin, Rajan Singh, Pierfrancesco Visaggi, and Frank Zerbib provided intellectual input and approved final manuscript. Maura Corsetti co-conceived of idea, provided intellectual input, and approved final manuscript. No competing interests declared.

Read the paper · More papers on PaperTik