A data-driven approach to cancer classification: the role of pre-processing in model accuracy and efficiency

Mukrimah Nawir, Amiza Amir, Nik Adilah Hanin Zahri, Masyitah Abu · 2025

This paper explores the impact of data pre-processing on the performance of machine learning (ML) models for cancer classification using medical images. The analysis focuses on the classification accuracy and computational efficiency of four transfer learning models: VGG16, DenseNet121, ResNet50, and MobileNetV2. The experiments are conducted on a large-scale breast cancer dataset at 400x magnification, consisting of 1,148 samples. The study reveals that pre-processing techniques significantly influence both the accuracy and the training time of the models. While the application of preprocessing consistently improves classification accuracy, particularly for DenseNet121 and ResNet50, It is also increases the training time, especially for the VGG16 model. These findings highlight the critical balance between pre-processing approaches and computational resources when optimizing deep learning models for medical image classification tasks. The results demonstrate that data-driven approaches, with careful consideration of pre-processing, can enhance cancer classification models’ effectiveness, albeit with increased training costs.

Read the paper · More papers on PaperTik