Toward clinical reliability: Visualizing and interpreting AI-based classification in peripheral blood smear analysis
Hiroaki Iwata, Tsukie Shibayama, Miku Watanabe, Hisashi Shimohiro · Machine Learning with Applications · 2025
• The developed AI model accurately classified peripheral blood smear cells. • Integrated Grad-CAM visualized crucial nuclear and cytoplasmic diagnostic features. • Predictions for early granulocyte precursors and neutrophils were interpretable. • Enhanced transparency addressed AI “black-box” issues in hematopathology workflows. • Findings support clinical adoption of explainable AI in hematologic diagnostics. With the advent of digital microscopy, the International Council for Standardization in Hematology recommends digital imaging and artificial intelligence (AI) algorithms for automatically classifying blood cells in peripheral blood smears to enhance diagnostic efficiency and accuracy. Nevertheless, while early AI studies have shown promising results in classifying white blood cells, the prediction process often remains unclear. Herein, we aimed to build a highly accurate model and visualize the basis of its predictions. The dataset comprised peripheral blood smear images of normal cells from individuals without infections, hematological disorders, or tumors, who were not undergoing any drug treatment at the time of blood collection. The images were obtained using a CellaVision DM96 analyzer at the Core Laboratory of Hospital Clínic de Barcelona. We used VGG16 and ResNet50 with transfer learning on ImageNet and applied the Grad-CAM method to visualize the image regions on which the model focused for classification. The model effectively recognized features, such as nuclear indentation and cytoplasmic color, which are crucial for classifying promyelocytes, myelocytes, and metamyelocytes. Traditionally, the basis of AI model predictions has been opaque, posing a challenge for medical applications. Our visualized classification basis clarifies the decision-making process of the model. These insights suggest that understanding these features can make the predictions of AI models more reliable and interpretable. Our findings improve diagnostic efficiency and suggest the potential of AI-based diagnostic support systems. Future research should validate this model’s performance using more extensive datasets and different cell types to enhance its reliability and practicality.