Statistical Performance Evaluation of the Deep Learning Architectures Over Body Fluid Cytology Images

Ersin Uysal · IEEE Access · 2025

The analysis of body fluid cytology images is a faster and easier diagnostic test than traditional methods for detecting cancer cells. Currently, there is limited statistical information on the performance of metric measurements obtained from different deep learning architectures (DL). In this study, the simulation results produced by different deep learning architectures are statistically analyzed in detail and the differences between the models are analyzed. Five different DL architectures (VGG16, VGG19, MobileNet, InceptionV3 and DenseNet121) are evaluated. The mean values of the metric measurement parameters (consistency, specificity, precision, accuracy and F1-score) obtained from the confusion matrix were evaluated using the non-parametric statistical methods Kruskal-Wallis and Mann-Whitney U tests. In addition, pairwise comparisons between AUC (Area Under the Curve) values obtained from ROC analysis were statistically evaluated. When the results were analyzed, the highest success performances were obtained in VGG16 with 98.60% sensitivity, InceptionV3 with 95.69% specificity, DenseNet121 with 95.02% accuracy, DenseNet121 with 89.48% F1-score and DenseNet121 with 91.42% AUC. The results obtained in the study showed that the criteria used in the evaluation metrics have different performances for different CNN architectures. In future studies, it will guide researchers on which CNN architecture gives strong results in which evaluation metric.

Read the paper · More papers on PaperTik