Multi-Modal Deep Learning for Malaria Diagnosis: Integrating CNNs and Vision Transformers for Enhanced Parasite Detection
P. Surendar, CH Hussaian Basha, Mohamed Suhail Mohamed Nabi, L. Kavitha, P. Vijayakumar, S Surender · 2025
Despite much progress towards malaria eradication, malaria remains a major global health problem, especially in areas lacking expertise in expert microscopists. Although automated malaria diagnosis by deep learning has recently emerged as a promising approach, current methods are mostly unsuitable for generalization and interpretability. To the best of our knowledge, in this study, we propose a multi-modal deep learning framework for parasite detection in microscopic blood smear images that combines Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). However, CNNs are powerful at extracting local spatial features while ViTs exploit self attention to generate a complete feature representation. We train and validate our model on publicly available malaria image datasets using state of the art data augmentation strategies, transfer learning, and state of the art attention mechanisms in order to maximize diagnostic accuracy. Experimental results show the superior performance of the proposed hybrid compared to conventional CNN based models on the basis of accuracy, precision and recall. In addition, explainability techniques like Grad-CAM and attention heatmaps explain what the model is deciding. The contribution of this study is to demonstrate the promise of multi modal deep learning in transforming malaria diagnosis to more interpretable, scalable, and AI based healthcare.