Posterior Probability of Discovery and Expected Rate of Discovery for Multiple Hypothesis Testing and High Throughput Assays
Yihua Zhao · Journal of the American Statistical Association · 2011
Technologies of measuring millions of quantities at once have been rapidly developed and used in biological and biomedical research and other fields in the past decade. Yet a key issue remains unsettled, namely, how to control spurious findings for the ensuing massive number of hypothesis tests. An emerging consensus is to control false discovery rate (FDR), with FDR defined as the expected proportion of true nulls among discoveries, and a discovery the rejection of a null hypothesis. However, the very concept of counting true nulls, implicitly or explicitly, is problematic, for nulls are rarely true in reality. We propose an approach that is philosophically different from the FDR and other approaches. Taking advantage of the massive measurements, we can and should directly evaluate the reproducibility of a discovery by calculating the posterior probability of discovery (PPD) given observed data. A discovery with a great PPD is deemed to be highly reproducible, that is, not spurious. For a subset of hypotheses tested, mean PPD yields the expected rate of discovery (ERD), a measure useful for various applications such as subset enrichment analysis. We present here the rationale, theoretical basis, and an algorithm for computing PPDs and ERDs from data. Their validity, utility, and optimal performance are demonstrated using both simulated and real data. Supplementary material is available online.