Reliable Medical Data Augmentation for Deep Learning: a Case Study on Breast Cancer Prediction
Samaké Adama, Diassana Fatoumata dite Soukoura, Ismaël Koné, Boulmane Lahsen · 2024
The limited availability of data poses a challenge for the effective application of deep learning techniques. To address this, various data augmentation methods have been developed. However, a significant limitation of these techniques is that they do not always ensure the integrity of the generated data. This means there is a risk of producing data that does not accurately reflect the original, potentially misrepresenting the phenomenon being studied. This drawback is particularly relevant when applying these augmentation techniques to medical images. In this paper, we propose an implicit data augmentation approach for classification problems, regardless of the data type. Starting with any classification problem, our approach first reformulates it into a binary classification task. Then, for this new task, we generate a new dataset composed of combinations of examples from the original dataset. As a result, the dataset size increases substantially, leading to implicit data augmentation. To solve this newly formulated classification problem, we define a generic neural network architecture. Preliminary results from a case study using the BreastMNIST dataset show promising im-provements. Specifically, we observed a significant performance increase due to our implicit data augmentation approach. A key advantage of our method is that it preserves the integrity of the training data, offering greater reliability for applying deep learning models to medical data.