Re: An Observational Study of Deep Learning and Automated Evaluation of Cervical Images for Cancer Screening

Robert G. Pretorius, Jerome Leslie Belinson · JNCI Journal of the National Cancer Institute · 2019

The manuscript by Hu et al. (1) concluded that automated visual evaluation of cervigrams was more accurate in diagnosing cervical intraepithelial neoplasia (CIN) 2, CIN 3, or cancer (CIN 2+) than were visual cervigram result (P < .001), conventional Pap smears (P < .001), liquid-based cytology (P < .001), and high-risk human papillomavirus (hrHPV) testing (P < .001). In Table 2 of the manuscript, the age-specific sensitivity of automated visual evaluation was reported as 82.1% to 97.7%. We believe that the sensitivity of automated visual evaluation of cervigrams for CIN 2+ has been inflated and that the relatively low accuracy of hrHPV tests for diagnosis of CIN 2+ as shown by the receiver operating characteristic curve in Figure 4 F is inconsistent with previous research. The sensitivity of automated visual evaluation of cervigrams for CIN 2+ is likely inflated because the gold standard for diagnosing CIN 2+ in this data set was colposcopic-directed biopsy of the worst lesion. Because cervigrams and colposcopy both use visual detection of lesions after application of acetic acid, it is likely that they are correlated (ie, they detect and miss the same subsets of CIN 2+). When the screening test and gold standard for diagnosing the disease are correlated, there is a mathematical relationship showing that the sensitivity of the screening test is inflated (2). In 2007, we showed that an abnormal visual inspection after application of acetic acid (VIA) was correlated with diagnosis of CIN 2+ by colposcopic-directed biopsy (kappa = .440); this correlation resulted in an inflation of sensitivity for CIN 2+ of VIA by 20% (2). We have reported receiver operating characteristic curves for diagnosis of CIN 3+ by hrHPV tests (3), which show much greater accuracy than that reported in Figure 4F of the Hu et al. article (1). We have reported that hrHPV tests had a sensitivity for CIN 2+ of 96.8% and a specificity of 79.7% (Youdens’ index of .765), and VIA had a sensitivity for CIN 2+ of 45.9% and a specificity of 92.2% (Youdens’ index of .381), clearly showing that hrHPV tests are more accurate than VIA (2). Because the methodology in the Hu et al. (1) manuscript results in inflation of sensitivity of automated visual evaluation of cervigrams and the low accuracy of hrHPV tests as shown in Figure 4F is not consistent with prior research, we question the authors’ conclusion that automated evaluation of cervigrams is more accurate in diagnosing CIN 2+ than is hrHPV testing. Neither Robert G. Pretorius nor Jerome L. Belinson has conflicts of interest to disclose related to this correspondence.

Read the paper · More papers on PaperTik