Comparison between CNN, YOLOv11, and vision transformers deep learning architecture for B-ALL detection
Muhammad Rafi Isnaen, Nathanael Mahardika Susilo, Chris Bernard, Kristien Margi Suryaningrum, Ricky Reynardo Siswanto · Procedia Computer Science · 2025
B-Cell Acute Lymphoblastic Leukemia (B-ALL) detection faces challenges due to subjective and time consuming traditional diagnostics. Majority of existing studies are primarily aimed at maximizing one model without any basis of reason. This study strictly compares three deep learning architectures—Convolutional Neural Network (CNN), You Only Live Once (YOLO), and Vision Transformers (ViT)—for multi-class classification of B-ALL subtypes (Benign, Early Pre-B, Pre-B, Pro-B) using standardized pipeline preprocessing, augmentation, and evaluation with the same settings. Comprehensive metrics (training/validation accuracy and loss, test accuracy, precision, recall, F1-Score, Macro/Weighted AUC, and categorical cross-entropy loss). CNN reveals superiority, achieving 98.0% test accuracy, stable graph training/validation, minimal cross-entropy loss (0.1009), and robust AUC scores (macro : 0.9974, weighted : 0.9978). While ViT showed 86.5% test accuracy, it exhibited data volume sensitivity (more fluctuations on train/validation accuracy and loss). Despite of YOLOv11 which supports classification, have high precision, recall, and F1-Score (0.99941, 0.99513, 0.99726 respectively), YOLOv11 have the lowest AUC scores and highest categorical cross-entropy loss (macro : 0.5139, weighted : 0.4679, loss : 11.0146 respectively). This study applies a rigorous benchmark for B-ALL classification diagnosis, offering a basis and more insight into the research allocation for a development focus to a model architecture to achieve medical application in further studies.