Using the swfdr package to estimate false discovery rates conditional on covariates
Simina Maria Boca, Tomasz Konopka, Leah R. Jager, Jeffrey T. Leek · 2020
We will present the swfdr R/Bioconductor package, which allows users to estimate the false discovery rate conditional on covariates. Multiple hypothesis testing is a common concern in modern science, including genome-wide association studies, transcriptomics, and proteomics. In many cases, tens or hundreds of thousands of statistical tests may be performed, leading to the necessity of multiple testing adjustments, often in terms of controlling or estimating the false discovery rate (FDR). Additional information, termed “meta-data,” “co-data,” or “feature-level covariates” – that may include allele frequencies, sample sizes, or gene or protein lengths – may be considered “prior information” and can be included in the decision on whether a particular null hypothesis should be rejected. We define an extension of the FDR that considers meta-data. Our approach for estimating the FDR conditional on covariates uses a regression framework to model the proportion of true null hypotheses. A recent independent benchmarking paper showed that this technique has excellent properties, including superior power and flexibility compared to other methods. Moreover, our approach allows for multivariate covariates, which some of the competing methods do not. We will introduce our method via an example where we consider FDR adjustments for sample size and allele frequency in a BMI genome-wide association study meta-analysis. Specifically, we will describe the lm_qvalue function, which estimates the q-values conditional on covariates.