Assessing the performance of risk prediction models

Dusko G. Nezic · European Journal of Cardio-Thoracic Surgery · 2020

I read with great interest the article by Smith et al. [1], which aimed to externally validate the postoperative atrial fibrillation (POAF) prediction model developed by Passman et al. [2] back in 2005. Two key aspects characterize the performance of a prediction model: calibration and discrimination. These should be reported in all papers reporting on prediction models [3]. During the external validation of a risk prediction model, we have to check its discrimination and calibration on an external patient cohort. Discrimination can be assessed by the area under the receiver operating characteristic (ROC) curve. The area under the ROC curve represents the percentage of randomly drawn pairs (a patient with an event paired with one without) in which the patient who had an event (in this case, POAF after major non-cardiac thoracic surgery) had a higher risk score than a patient without. In other words, discrimination differentiates low-risk from high-risk patients. For a binary outcome, c-index is identical to the area under the ROC curve, which plots the sensitivity (true positive rate) against 1—(false positive rate) for consecutive cut-offs for the probability of an outcome. The discriminative power is thought to be excellent if the area under the ROC curve is >0.80, very good if >0.75 and good (acceptable) if >0.70 [3]. Unfortunately, we could not find any discriminatory data throughout the manuscript [1]. Furthermore, 3 cited manuscripts, presenting prediction models for POAF demonstrated poor C-statistics (0.62—Rao et al. [4]; 0.67—Onaitis et al. [5]; 0.65–0.73 for different surgical groups—Passman et al. [2]). If the discrimination of a model is good, but the calibration is not, the model can be made more accurate by recalibration. However, the converse is not true [6]. If discrimination is not good, as is the case for almost all the cited prediction models [2, 4, 5], none of these risk prediction models can be used as a predictor of POAF onset. Calibration refers to the agreement between observed number of events and predicted probability of occurrence of these events. Calibration is preferably reported graphically [3] with predicted outcome probabilities (on the x-axis) plotted against observed outcome frequencies (on the y-axis). The goodness-of-fit test, usually the Hosmer–Lemeshow test, measures the differences between observed and expected outcomes over deciles (10 groups of patients) of risk. A well-calibrated model gives corresponding P-value >0.05. Another useful statistical tool is the observed to expected event ratio. Ideally, this ratio equals one (the observed number of events equals expected occurrence of these events, thus indicating that the predictive model is perfectly calibrated). A value above one means that model underestimates occurrence of event, a value below one means that model overestimates occurrence of event. If the 95% confidence interval of the observed to expected event ratio includes the value of 1.0, the model is well-calibrated [3]. Both tools could be graphically presented using calibration plots [3].

Read the paper · More papers on PaperTik