A Multimodal and Interpretable Deep Learning Framework for Early Detection of Blood Cancer Using Clinical Data
G. Chinna Pullaiah · Communications on Applied Nonlinear Analysis · 2025
Diagnosing blood cancer at early stages presents challenges due to limitations in methods that rely on single data sources, which often fail to provide comprehensive insights. The study introduces a Multimodal and Interpretable Deep Learning Framework (MIDLF) that combines patient symptoms, imaging, and laboratory results to predict blood cancer with explainable outputs. The framework integrates a graph-based model for encoding symptoms, a hybrid Vision Transformer (ViT) and Convolutional Neural Network (CNN) for imaging analysis, and an attention-driven module for laboratory data. A gated attention mechanism fuses these inputs, capturing relationships across modalities. Preprocessing includes graph-based embeddings for symptoms, self-supervised learning for unlabeled imaging data, and generative models to address missing laboratory values. Evaluation used the BIOBANK dataset, which includes multimodal clinical data. The framework achieved an accuracy of 94.8%, a recall of 95.2%, and an AUC-ROC of 97.1%. These results were benchmarked against Cell Scoring Neural Network (CSNN) and Three-Stage Feature Selection with Twice-Competitional Ensemble Learning Method (TSFS-TCEM), demonstrating improved prediction consistency across metrics. Clinician reviews validated feature importance outputs from SHAP and Layer-wise Relevance Propagation, showing alignment with clinical reasoning. This framework addresses challenges in integrating diverse clinical data and explaining model predictions. Results suggest that combining symptoms, imaging, and laboratory data enables more informed diagnostic predictions, facilitating early detection of blood cancer.