Artificial Intelligence and radiologist interpretation of screening mammography: Classification and comparison of challenges with strategies for difficult cases
Zhengqiang Jiang, Ziba Gandomkar, Phuong Dung Trieu, Seyedamir Tavakoli Taba, Melissa L. Barron, Sarah Jayne Lewis · European Journal of Radiology Artificial Intelligence · 2025
The relationship between malignancy probability generated by an Artificial Intelligence (AI) model and the correct classification rate (CCR) by radiologists reading screening mammograms is explored in this paper. This study explores if radiomics features are statistically difference between different groups of easy, intermediate, and difficult-to-interpret cases. The Globally-aware Multiple Instance Classifier (GMIC) AI model with transfer learning on an Australian screening mammogram dataset with 1712 cases (856 malignant and 856 normal) was used in this study ( Sydney-GMIC). Different groups of easy, intermediate, and difficult-to-interpret cases were generated based on different ranges of malignancy probability and CCR. Radiomics features of mamography cases were computed on mamography images and significant tests were conducted on 14 radiomics features between different groups of cases. At least 73 Australian radiologists and the Sydney-GMIC read a second Australian screening mammogram dataset of 540 cases (179 malignant and 361 normal). For cancer cases, the majority were classified as ‘easy and intermediate’ cases for both Sydney-GMIC and radiologists. Of the easy and intermediate cases, AI attributed 52 % cases to easy category and 48 % cases to intermediate group, whereas the radiologists CCR attributed 60 % to easy and 40 % to intermediate group. For normal cases, most cases were ‘easy’ for the Sydney GMIC and ‘easy’ and ‘intermediate’ cases for radiologists. Near perfect agreements were found between both easy and intermediate cases and the contrast radiomics feature was significantly different for these two groups. 82 out of 140 Radiomics feature comparisons (e.g., contrast, correlation, coarseness and complexity with p-values <0.05 between a group of easy for AI and radiologists, and another group of difficult for AI and radiologists) were statistically significant difference between two different groups with normal cases. The number of cases that were both difficult and easy for radiologists and AI were similar, indicating good concordance. AI models could show promise if deployed as a supportive reader for low to intermediate risk cases, along with the use of radiomics for both AI training and the complement of AI learned features. Double-reader strategies involving AI and human readers may yield improved performance via correct classification rates for malignant cases. • Study matches malignant probability score by AI and correct classification rate by radiologists reading screening mammograms. • Explores why cases that are difficult or easy for radiologists to interpret pose similar challenges or ease for AI models. • Investigates which radiomic features of mammographic cases are significant in determining case difficulty.