Conflict avoidance in mammography: filtering datasets for breast cancer risk prediction

Alistair Taylor-Sweet, Adam Perrett, Stepan Romanov, Raja Ebsim, D. Gareth Evans, Susan Astley · 2025

Deep learning methods can provide reliable estimates of breast cancer risk which could be used to personalise screening or identify women for whom preventive interventions would be appropriate. Some breast cancer risk prediction methods incorporate percent mammographic density estimated by expert readers and recorded on Visual Analogue Scales (VAS). VAS can be used to train deep learning models, but one of the main challenges facing this area of research is a lack of reliable labels; because of inter-reader variability, it is preferable to take an average density assessment across two expert readers. Where data from two readers are available, the removal of examples where readers disagree could be used to filter the data down to increase the quality of training data for deep learning models and hence improve accuracy. Using the Predicting Risk Of Cancer At Screening (PROCAS) dataset of mammograms from 38,861 women with unprocessed GE images and ResNet50 to predict mammographic density, several models were trained to compare how filtering the dataset using disagreement between experts affects accuracy with thresholds on disagreement between experts of 1, 5, 10, 15, 25 and 30%. The Odds Ratio (OR) between the risk of cancer in the lowest and highest quintiles of predicted densities is used to compare performance using an unfiltered testing set. Results show that there is a slight improvement when filtering the dataset to exclude examples where experts disagree by more than 25%, from an OR of 3.054 (95% CI 2.81-3.3) in the model trained on unfiltered data to 3.204 (95% CI 2.96-3.45) for the filtered model.

Read the paper · More papers on PaperTik