MammoClean: Toward Reproducible and Bias-Aware AI in Mammography Through Dataset Harmonization
Yalda Zafari-Ghadim, Hongyi Pan, Görkem Durak, Ulaş Bağcı, Essam A. Rashed, Mohamed A. Mabrok · IEEE Access · 2026
The development of clinically reliable artificial intelligence (AI) systems for mammography is hindered by profound heterogeneity in data quality, metadata standards, and population distributions across public datasets. This heterogeneity introduces dataset-specific biases that can compromise the generalizability of AI models, a fundamental barrier to clinical deployment. We present MammoClean, a public framework for dataset standardization and bias identification in mammography datasets. MammoClean standardizes case selection, image processing (including laterality and intensity correction), and unifies metadata into a consistent multi-view structure. We provide a targeted review of breast anatomy, imaging characteristics, and public mammography datasets to systematically identify key sources of heterogeneity. We apply MammoClean to three large-scale heterogeneous datasets with detailed annotations (CBIS-DDSM, TOMPEI-CMMD, VinDr-Mammo) as a focused case study to quantify distributional shifts and illustrate the impact of data corruption on AI model performance. Applying MammoClean to the three datasets, we quantify substantial distributional shifts in breast density and abnormality prevalence. We further illustrate the direct impact of data corruption: AI models trained on corrupted datasets exhibit notable performance degradation compared to their curated counterparts. By using MammoClean to identify and document sources of heterogeneity, researchers can more reliably prepare unified multi-dataset training corpora as a foundation for developing robust models. MammoClean provides a reproducible pipeline for bias-aware dataset preparation in mammography, facilitating fairer comparisons across studies. The open-source code will be publicly available from: https://github.com/Minds-R-Lab/MammoClean