Hybrid Diffusion Architecture for High-Fidelity Medical Image Synthesis

Jiarui Wang, Lu Chen · 2025

Medical image synthesis remains a critical challenge in biomedical engineering due to the scarcity of annotated datasets and the computational complexity of generating high-fidelity anatomical representations. This paper introduces a novel diffusion model framework designed to address these challenges through synergistic architectural innovation and dynamic parameter adaptation. Leveraging the strengths of transformer-based global context modeling and U-Net's hierarchical feature propagation, we propose the U-NetTransformer architecture—a hybrid backbone that replaces conventional convolutional modules with multi-head self-attention (MHSA) layers augmented by cross-scale feature interaction units. This design enables adaptive decoupling of structural generation and textural refinement processes, significantly improving anatomical coherence in synthesized images. To resolve inherent parameter-sharing conflicts in traditional diffusion models, we introduce the Temporal Conditional Mask Generator (TCMG), a lightweight module that dynamically activates task-specific sub-networks using learnable binary masks without requiring full model retraining. The proposed framework adopts a parameter-efficient co-optimization strategy, maintaining frozen pre-trained weights while training only mask generators and attention modules. Experimental evaluations on multi-organ CT/MRI datasets demonstrate substantial performance gains over baseline diffusion models. Clinical validation by board-certified radiologists confirms the pathological consistency and diagnostic usability of synthesized images, particularly in rare disease scenarios.

Read the paper · More papers on PaperTik