Model Structure Adaptability and Performance Analysis of Medical Image Denoising Under Multimodal Data

Linfei Xiao · 2025

Medical image noise originates from low-dose scanning protocols and differences in tissue density, which directly affect the accuracy of quantitative analysis of lesions. Current medical denoising research mainly uses traditional filtering and deep networks. The former is prone to over-smoothing, while the latter improves the effect but faces a bottleneck of cross-modal generalization. Given the performance differentiation problem of existing models in complex medical noise scenarios, this study systematically compares the adaptability of three typical architectures: Attention UNet, GAN, and PatchGAN under multilevel noise. Four types of noise models (low/high Gaussian, pepper and salt, mixed) are constructed based on self-built MRI/X-ray/CT multimodal datasets, and the correlation mechanism between PSNR/SSIM indicators and model structure is analyzed through experiments. Attention UNet performs best in low Gaussian noise, pepper and salt noise, and mixed noise scenarios, and its multiscale feature fusion and dynamic attention mechanism significantly improve noise localization accuracy; standard GAN has only a slight advantage in high Gaussian noise SSIM, but its shallow architecture leads to generally low PSNR; PatchGAN has an essential conflict between adversarial training objectives and pixel-level reconstruction, and its PSNR/SSIM is significantly inferior to the baseline model in most scenarios. UNet-type models should be used as the preferred architecture for medical denoising because their encoding and decoding structure is highly compatible with the characteristics of medical images; the adversarial training paradigm should be used with caution in diagnostic scenarios that require strict fidelity, and its structural characteristics are more suitable for prior learning of specific noise distributions.

Read the paper · More papers on PaperTik