Multimodal Federated Learning for Healthcare: Adaptive and Private Fusion

Yiheng Wang · Applied and Computational Engineering · 2025

Federated learning in medical image analysis has emerged as a critical paradigm to address data privacy and cross-institutional collaboration challenges. This study proposes a privacy-preserving multimodal federated learning framework that integrates Residual Network (ResNet) 50- Feature Pyramid Network (FPN) architectures with cross-modal attention mechanisms to enhance diagnostic accuracy while safeguarding sensitive patient data. Specifically, the framework employs a dual-stream design: an FPN extracts hierarchical visual features from medical images (e.g., dermoscopic lesions, lung opacities), while a transformer-based module aligns these features with clinical text embeddings (e.g., lab reports, patient histories). Local models are trained on decentralized datasets using hybrid losses combining cross-entropy and contrastive alignment, with differential privacy (σ=0.5) applied to gradients to ensure (ε=1.5, δ=10−5) compliance. Experiments on the International Skin Imaging Collaboration (ISIC) 2019 and Novel Coronavirus Pneumonia (COVID-19) radiography datasets demonstrate the framework’s superiority, achieving 92.4% classification accuracy (vs. 85.7% for single-modal federated learning) and robust performance under non-IID data distributions. The results highlight the efficacy of multi-scale feature fusion and privacy-aware aggregation in medical Artificial Intelligence (AI), offering a scalable solution for collaborative diagnosis without compromising data sovereignty. This work advances ethical AI deployment in healthcare, enabling institutions to leverage heterogeneous multimodal data while adhering to regulatory standards like the Health Insurance Portability and Accountability Act (HIPAA) and the General Data Protection Regulation (GDPR).

Read the paper · More papers on PaperTik