C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation
Jeonghyeok Do, Jaehyup Lee, Seungchul Lee, Munchurl Kim · IEEE Transactions on Circuits and Systems for Video Technology · 2026
Synthetic Aperture Radar (SAR) imagery provides robust environmental and temporal coverage (e.g., during clouds, seasons, day-night cycles), yet its noise and unique structural patterns pose interpretation challenges, especially for non-experts. SAR-to-EO (Electro-Optical) image translation (SET) has emerged to make SAR images more perceptually interpretable. However, traditional approaches, trained from scratch on limited SAR-EO datasets, are prone to overfitting and struggle with dataset inconsistencies. To address these challenges, we introduce Confidence Diffusion for SAR-to-EO Translation (C-DiffSET), a framework leveraging a pretrained Latent Diffusion Model (LDM) to effectively adapt its extensive generative priors from natural images to the EO domain. Our investigation reveals that the pretrained VAE encoder effectively aligns SAR and EO images within a shared latent space, demonstrating robustness even to varying noise levels in SAR inputs. To further improve pixel-wise fidelity for SET and mitigate artifacts from temporal discrepancies, such as appearing or disappearing objects, we propose a novel confidence-guided diffusion (C-Diff) loss. This loss dynamically guides the diffusion process to down-weight penalties in uncertain regions, thereby enhancing structural accuracy. C-DiffSET achieves state-of-the-art (SOTA) results on multiple benchmark datasets (QXS-SAROPT, SAR2Opt, SpaceNet6, Stellar-Vision), significantly outperforming recent image-to-image translation methods and specialized SET methods across all standard metrics. The source code and trained models are available at https://github.com/KAIST-VICLab/C-DiffSET.