Multi-modal fusion and transferable deep learning for rare disease detection: a CNN-Transformer framework with cross-domain adaptation on limited CT and MRI data

Jingyu Tang · Advances in Engineering Innovation · 2025

Medical imaging diagnosis of rare diseases faces the dual challenge of scarce labeled data and significant differences in equipment. This study proposes a hybrid framework integrating CT and MRI modalities, combining CNN and Transformer architectures and introducing an adversarial domain adaptation mechanism. The dedicated CNN encoder extracts fine-grained local features, while the Transformer module captures long-range cross-modal dependencies. The gradient inversion domain discriminator aligns the feature distribution of different scanning devices to ensure the device independence of the model. On the two public datasets of neurological diseases, intracranial hemorrhage and demyelinating lesions, the average accuracy rate of this model reached 91.3%, and the F1 score was 0.89, which is 5 to 10 percentage points higher than the single-mode and pure Transformer baseline. Ablation experiments confirmed that the domain-based adversarial transformer component and training contributed to significant performance gains. In the cross-domain (CT to MRI) experiment, the domain adaptation technique increased the F1 score from 0.74 to 0.84. These results highlight the effectiveness of local feature extraction, global context modeling, and collaborative adversarial alignment schemes in multi-agency scenarios with sparse data. Future research will be extended to three-dimensional volume data, integrate semi-supervised learning of unlabeled images, and optimize the reasoning process of real-time clinical decision support.

Read the paper · More papers on PaperTik