Computer‐aided detection should be used routinely to assist screening mammogram interpretation
Robert M. Nishikawa, Joshua J. Fenton, Colin G. Orton · Medical Physics · 2012
Computer-aided detection (CADe) is used for the interpretation of screening mammograms in many institutions, especially in the United States. Some suggest that the time is ripe to use it routinely for all mammographic screening programs worldwide. Others, however, question its routine use because they claim that there is no convincing evidence that it has any mortality benefit. This is the topic debated in this month's Point/Counterpoint. Let me begin my argument with the underlying assumption that breast cancer screening with mammography is effective at all six levels of effectiveness as outlined by Fryback and Thornbury.1 Mammographic screening is effective because it can detect breast cancers early enough for them to be treated effectively. Computer-aided detection can help radiologists to reduce their cancer miss rate and thereby promote earlier detection. The DMIST study showed that ∼50% of women who have breast cancer have their mammograms read as normal.2 Computer-aided detection could, therefore, have a large impact on the false-negative rate of mammography. In a Point/Counterpoint debate published in 2006, I argued that the effectiveness of CADe in screening mammography was unproven.3 Since then, several additional clinical studies have been published but, more importantly, our understanding of CADe and how to measure its effectiveness has improved. Compared with 2006, research on the effectiveness of CADe, while still not definitive, clearly shows that its use can reduce radiologists’ miss rate but, on the negative side, it increases the radiologists’ recall rate. However, the increase in number of cancers detected (∼10%) is comparable to the increase in recall rate (∼12%).4 I believe that this recall rate is slightly elevated because radiologists are still learning how to the use the system effectively. The ratio of the number of cancers detected to the number of recalls is the positive predictive value of screening (PPV1). If PPV1 is essentially unchanged—a 2% decrease (1.10/1.12)—and screening mammography is effective, then screening mammography with CADe must also be effective, since the number of recalls for every cancer detected is the same whether CADe is used or not. Further, since mammography is “cost effective” at this PPV1, finding more cancers at a fixed PPV1 must increase the benefit of screening when CADe is used. Perhaps the strongest evidence for CADe allowing earlier detection of cancer is from the two studies published by Fenton et al.5,6 In these studies, the fraction of cancers that were detected as DCIS increased when CADe was used, with a corresponding decrease in Stage 1 invasive cancers. Note that all cancers a radiologist detects because of using CADe are mammographically visible. This means that the cancer will eventually be detected mammographically—unless it is detected as an interval cancer (a cancer that is detected nonmammographically between screening exams)—whether or not the radiologist uses CADe.7 Therefore, CADe is unambiguously detecting a cancer at an earlier stage. While CADe is an effective tool for screening mammography, an increase in sensitivity of 10%, in light of a 50% miss rate when CADe is not used, is moderate at best. New approaches to implementing CADe clinically and better training of radiologists in how to use CADe, however, can potentially enhance its benefit. Breast cancer screening is performed among healthy, asymptomatic women with the goal of reducing their risk of dying from breast cancer. Before recommending routine application of CADe during screening mammography, one must be confident that its use in mammography is helping, not hurting, women. While consideration of patient harms and costs are important, fundamentally we must be confident that CADe augments screening mammography's impact on breast cancer mortality. Within this framework, one cannot soundly recommend routine use of CADe during screening mammography. There are no data showing that CADe use during screening mammography is associated with reduced breast cancer mortality, and analyses of national mammography data suggest that CADe use is associated with little if any change in mammography performance or in the stage or size of detected breast cancers.6 Under optimal conditions,8 CADe may incrementally increase the detection of noninvasive breast cancers (ductal carcinomas in situ). But mammography's mortality benefit largely derives from early detection of invasive breast cancers,9 so even optimal CADe use is unlikely to reduce breast cancer mortality risk.10 Indeed, finding more ductal carcinoma in situ of ambiguous lethal potential is a potential harm of mammography, as many in situ carcinomas may be overdiagnosed and unnecessarily treated, particularly in older women. Computer-aided diagnosis also clearly increases the false-positive rate of screening mammography. Leading to considerable anxiety in most women, false-positive mammography constitutes a substantial aggregate harm of CADe when multiplied by the millions of women now exposed to CADe in U.S. practice.11 Routine CADe use also comes with significant monetary costs. Although CADe is already widely used in the United States,11 if CADe use were extended to all of the 31 × 106 screening mammograms performed/year in the United States, CADe-attributable costs from added insurance payments (∼$12/mammogram) and diagnostic testing after additional false-positive mammograms (∼$450/false-positive) would exceed $550 × 106 annually.5 One might argue that mammography is poorly reimbursed, and CADe fees represent an essential revenue stream for mammography practices. But Congress did not add the coverage for CADe to the Medicare benefit as a means of subsidizing mammography practice. Rather, CADe coverage was the result of adroit industry lobbying and a Congress that was poorly prepared to evaluate the limited evidence of CADe's effectiveness.12 Reimbursement incentives explain the broad and rapid adoption of CADe in the United States as compared to other developed countries, and Congress would be wise to rescind Medicare coverage for CADe, removing the incentive for its routine use in the United States. Some may argue that theoretical benefits of routine CADe use justify its harms and costs. Similar arguments were made in support of routine self-breast examination until large randomized trials demonstrated harm without benefit.13 An analogous adverse risk–benefit ratio may be the ultimate truth with CADe. While large randomized trials of CADe use may be impractical, CADe should ideally be used only in the context of research protocols and certainly not routinely, as is regrettably the case in U.S. practice. I concede to my colleague that the broad adoption of CADe in the United States is a result of reimbursement, and the decision to approve reimbursement was not based upon clinical evidence demonstrating its effectiveness. Nonetheless, this does not mean that CADe is ineffective and should not be used routinely. Let me address two points from Dr. Fenton's opening statement. Dr. Fenton argues that CADe should not be used routinely because no study has shown that its use can reduce breast cancer mortality. Given that argument, screening with breast MRI or digital mammography should not be performed. In fact, conventional screening mammography was common long before there was consensus on a mortality benefit, which occurred 20 years after it became used routinely. While there is merit to Dr. Fenton's argument, the end result would be a stifling of innovation and many new imaging techniques would never see the light of day. To study the impact of any screening technology on breast cancer mortality requires tens of thousands of women, tens of millions of dollars, and over 10 years of follow-up. Clearly, some figure of merit other than mortality needs to be used as a criterion for clinical implementation. One possible surrogate end point could be cancer stage. Dr. Fenton argues that the only change in cancer characteristics is an increase in the detection of DCIS with a reduction in Stage I cancers. As I pointed out in my opening statement, CADe will not lead to overdiagnosis, because all cancers detected by CADe are mammographically visible and will be detected eventually without CADe. Therefore, based on Dr. Fenton's data, CADe is clearly leading to earlier detection. Since cancer stage alone does not predict mortality, I believe further clinical studies are needed to evaluate the effectiveness of CADe. These studies can only be done if CADe is being widely used clinically. Dr. Nishikawa implies that our work with the Breast Cancer Surveillance Consortium (BCSC; Ref. 6) demonstrated a favorable shift in breast cancer stage with CADe. In unadjusted analyses, CADe use was associated with a reduced rate of detection of invasive breast cancer but no change in the rate of detection of DCIS. In adjusted analyses, however, CADe was associated with a nonsignificant trend toward increased detection of DCIS but no differences in rate of detection of invasive cancer or in the diagnosis of early stage invasive cancers. As discussed in my opening statement, any increase in the detection of DCIS is of uncertain clinical significance in light of the ambiguous relationship between early DCIS detection and reduced breast cancer mortality.9 More importantly, it is misleading to state that “CADe is unambiguously detecting a cancer at an earlier stage” when adjusted analyses suggested little, if any, impact of CADe on breast cancer stage. Citing his own work,4 Dr. Nishikawa states that CADe reduces the “miss rate” of screening mammography by 10%, implying that CADe improves the sensitivity of screening mammography. However, BCSC data and a meta-analysis suggest little if any impact of CADe on screening sensitivity or breast cancer detection rates.6,14 Similarly, Dr. Nishikawa states that the PPV1 is “essentially unchanged” with CADe.4 On this basis, he reasons that CADe must be effective, given that screening mammography is also effective. The preponderance of other evidence, however, suggests that CADe is associated with reduced PPV1 (about a 16% relative decline from 4.3% to 3.6% in the BCSC data), increased recall rates and little, if any, improvement in cancer detection rates.6,14 The effectiveness of CADe should ultimately be based on whether it reduces breast cancer mortality in community practice. Evidence to date suggests that a breast cancer mortality benefit from current CADe technology is highly unlikely.10 Absent a mortality benefit, routine use of CADe during screening mammography cannot currently be justified.