The Problem With Predictive Values: Are We Using the Right Metrics for Preclinical Prediction of Drug Hepatotoxicity?
Mitchell R. McGill · Toxicological Sciences · 2018
Sensitivity, specificity, and predictive values are important to understand the real-world utility of biomarkers in serum, urine, body fluids, and tissue. Recently, however, the concept has been applied to prediction of drug hepatotoxicity during pre-clinical testing by Vorrink et al. (2018) and others. These colleagues have used predictive values as a way of expressing the ability of novel in vitro model systems to predict drug-induced liver injury (DILI), and as metrics that can be used to compare those systems. However, use of predictive values in that way is problematic. Before going further, we must define these terms. The sensitivity of a test for a condition, whether the test is a biomarker for diagnosis or an in vitro experiment to predict DILI, is the percentage of the condition-positive population that will have a positive test result. For example, if a test for lung cancer is positive in 80% of lung cancer patients, then the sensitivity of the test is 80%. Or if we have a selection of only drugs that are known to cause DILI and 60% of them cause toxicity in a particular in vitro liver model, then the sensitivity of the model for DILI is 60%. Conversely, specificity is the percentage of the condition-negative population that will have a negative test result. On the other hand, positive predictive value (PPV) is the percentage of the test-positive population that has the condition. That is, if you are positive for a test with PPV of 80% for a condition, then your probability of having the condition is 80%. Conversely, negative predictive value (NPV) is the percentage of the test-negative population that has the condition. Although sensitivity and specificity are helpful under the right conditions, they present 2 problems. First, they are not intuitive. If you have a test result and you only know sensitivity and specificity, you will likely have to think for a moment about the meaning of the result. Second, they do not tell the whole story. Suppose we want to compare a biomarker for cardiovascular disease (CVD) with one for chronic obstructive pulmonary disease (COPD) to determine which is more useful for its purpose. We could compare their values for sensitivity and specificity. If the CVD biomarker has sensitivity and specificity of 65 and 90%, respectively, while the COPD biomarker has values of 85 and 90%, then we might conclude that the COPD biomarker is better. However, that would be incorrect, as we can see from predictive values. Predictive values are an improvement over sensitivity and specificity because they take into account condition prevalence in addition to sensitivity and specificity. The prevalence of CVD in the adult US population is approximately 40%, while the prevalence of COPD is approximately 4%. A simple 2 × 2 table reveals that the PPV for the CVD biomarker is 81% while the PPV for the COPD biomarker is a mere 26%. Instructions for performance of these calculations are available in many textbooks (Deacon, 2009). Knowing only the above, it might be tempting to assume that predictive values are always better. However, now that it is clear that they depend upon prevalence, it should be obvious that the predictive values reported in studies using a selection (or “population”) of drugs that cause DILI will depend upon the percentage (or “prevalence”) of DILI-positive drugs in that selection. A study using a population of drugs in which 40% are known to cause DILI will yield better predictive values than one which in which only 4% are known to cause it, usually even if sensitivity and specificity are worse. There are 2 possible solutions. First, we can continue to use PPV and NPV, but make it standard to use a selection of drugs with the same percentage of DILI-positive agents (eg, 10%). In fact, a group of DILI experts could agree on a panel of drugs that everyone should use. Second, we can forego reporting of PPV and NPV in favor of metrics that are normalized for the known approximate prevalence of hepatotoxicity among new drugs entering clinical trials. Posttest probabilities can be calculated using likelihood ratios (easily determined from sensitivity and specificity) and a known prevalence (also called pretest probability) and give estimates of real-world predictive value. For example, if we assume that 10% of drugs entering clinical trials fail due to hepatotoxicity and we use data from Vorrink et al. (2018), we can calculate that the positive posttest probabilities for DILI would be 100% for both the 3D culture model used by Vorrink et al. (2018) and for the 2D model used in the study by Xu et al. (2008), while the negative posttest probabilities would be 3% and 5%, respectively. That is, when toxicity is observed in either system, the probability that the drug will cause DILI is 100%, while it is only 3%–5% when toxicity is not observed. Thus, the posttest probabilities have strikingly revealed that the 2D and 3D models are equivalent for prediction of DILI. (Of course, in reality, 100% is unachievable, and likely due to the small drug population; also, 100% specificity should be adjusted down to 99.9% for calculations.) Furthermore, the 3D models used in the other studies cited by Vorrink et al. (2018) actually have worse posttest probabilities. Of course, another issue with these studies is lack of standardized endpoints. Although some studies measure viability based on ATP, others measure reactive oxygen species (ROS), enzyme release, propidium iodide staining, or other endpoints. The use of different endpoints likely affects the above metrics. Some toxicants may cause transient ROS production or ATP depletion without killing the cell, so those end points would impart greater sensitivity, while cell death measures would yield greater specificity. Overall, there is a need for standardization of testing procedures and improved understanding of “predictive” statistics to facilitate comparison of in vitro model systems. I would like to acknowledge support from the American Association for the Study of Liver Diseases Pinnacle Research Award and from the University of Arkansas for Medical Sciences. Supplementary data are available at Toxicological Sciences online.