Letter to the editor on “Automated classification of fat-infiltrated axillary lymph nodes on screening mammograms”

Takuma Usuzaki, Kengo Takahashi, Ryusei Inamori · British Journal of Radiology · 2023

Dear Editor, We read with great interest a recent article “Automated classification of fat-infiltrated axillary lymph nodes on screening mammograms” by Song et al published in the British Journal of Radiology.1 Song et al developed an end-to-end deep learning model from image preprocessing to classifying the fat-infiltrated axillary lymph node (LN) using full-field digital screening mammograms. The parameters of the deep learning model were tuned by training (266 fat-infiltrated axillary LNs, 1558 normal LNs) and development (40 fat-infiltrated axillary LNs, 266 normal LNs) datasets. The tuned model was tested and validated using internal test (110 fat-infiltrated axillary LNs, 801 normal LNs) and external test (35 patients with fat-infiltrated axillary LN, 35 patients with normal LN) datasets, respectively. For the internal test dataset, the deep learning model achieved an accuracy of 0.99 and an area under the curve of receiver operating characteristics (AUC-ROC) of 1.00, respectively. This performance was validated with an accuracy of 0.82 and an AUC-ROC of 0.87 for the external test dataset. This outstanding performance achieved by Song et al indicates that deep learning is a promising approach to automatically detect fat-infiltrated axillary LNs in full-field digital mammograms. As the authors mentioned, these results have a clinical impact because fat-infiltrated axillary LNs associated with obesity,2 type 2 diabetes mellitus,3 and metastasis of breast cancer.4 However, there is a limitation the authors overlooked: the metrics were inadequately evaluated. Especially, the accuracy was overestimated for the internal test dataset due to its imbalance. Since the internal test dataset consisted of 110 fat-infiltrated axillary LNs and 801 normal LNs, at least, an accuracy of 801/911 = 0.879 can be achieved. In general, a deep learning model can achieve, at least, an accuracy that is equal to the proportion of positive samples in a dataset. When the accuracy of a deep learning model is evaluated for an imbalanced dataset, we need to distinguish true accuracy from apparent accuracy. We simulated the apparent accuracy by changing a true accuracy and the proportion of positive samples in an imbalanced dataset. In this simulation, we assumed that a positive case was predicted with the true accuracy. The apparent accuracy was repeatedly calculated by changing each true accuracy and proportion of positive samples from 0.50 to 1.00 in increments of 0.001. Figure 1 shows the relationships among an apparent accuracy, a true accuracy, and the proportion of positive samples in an imbalanced dataset. From this simulation, we can estimate the true accuracy for the internal test dataset Song et al analysed: when the apparent accuracy and proportion of positive samples in the internal test dataset was 0.990 and 0.879, respectively, the true accuracy was estimated as 0.932. The accuracy was overestimated by 0.990−0.932 = 0.058. This estimation was also shown in Figure 1. The result of the simulation. The heatmap shows the simulated apparent accuracies. The horizontal and vertical axes show the true accuracy and the proportion of positive samples in an imbalanced dataset, respectively. Using this figure, the true accuracy can be determined by an apparent accuracy and the proportion of positive samples in an imbalanced dataset. As an example, the solid circle, horizontal, and vertical broken lines represent the apparent accuracy (0.990), the proportion of positive samples (0.879), and the true accuracy (0.932), respectively. As shown by our simulation, the performance of a deep learning model can be overestimated for an imbalanced dataset. In the case of utilizing an imbalanced dataset, there are several ways to analyse the performance of a deep learning model. First, true accuracy is estimated from apparent accuracy and proportion of positive samples, as we demonstrated. We can similarly evaluate other metrics such as sensitivity, specificity, positive predictive value, and negative predictive value. Secondly, resampling techniques are used to rebalance the dataset.5 Resampling techniques fall into three groups depending on the method used to balance the class distribution: over-sampling,6 under-sampling,7 and hybrid-sampling.5 Song et al performed under-sampling in constructing the external test dataset. As an application of resampling techniques, cross-validation is often used in an imbalanced dataset.8 Thirdly, metrics with less susceptibility to imbalance such as F1 score and area under the curve of the precision-recall (AUC-PR) are used. These two metrics are defined using precision (positive predictive value) and recall (sensitivity) and do not have information on true negative. PR curve is specifically tailored for the detection of rare events and is more informative than AUC-ROC for imbalanced data.9 Song et al adequately used AUC-ROC for the external test dataset which was balanced, whilst AUC-PR should be used for the internal test dataset. In conclusion, adequate metrics need to be selected and evaluated for deep learning when an imbalanced dataset is used. We do not argue against using an imbalanced dataset for deep learning because class imbalance is unavoidable in most medical researches.5 We deeply appreciate the authors’ contribution to this field and thank them for providing an opportunity to discuss their paper. Takuma Usuzaki (Conceptualization, Investigation, Writing—Original Draft, Project administration), Kengo Takahashi (Conceptualization, Investigation, Writing—Original Draft, Project administration), and Ryusei Inamori (Investigation, Writing—Review & Editing) T. Usuzaki and K. Takahashi contributed equally to this work as first authors. This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. None declared.

Read the paper · More papers on PaperTik