SAMDiff-TTR: SAM2-Augmented Diffusion with Test-time Refinement for Small-Sample Infrared-to-Visible Image Generation
L. Y. Tian, Q. Shen, Y. Cao, D. Z. Sun, Z.L. Deng · Advanced Electromagnetics · 2026
Infrared-to-visible image generation provides an effective solution for augmenting visible-spectrum samples in small-sample scenarios and offers important support for cross-spectral information representation in electromagnetic sensing systems, thereby benefiting downstream visual perception tasks such as object detection in autonomous navigation and unmanned aerial vehicle (UAV) applications. However, existing methods often suffer from insufficient semantic consistency and weak generalization when only limited paired cross-modal data are available. To address this issue, we propose SAMDiff-TTR, a novel infrared-to-visible image generation framework that combines a diffusion-based generative model with semantic guidance from a pre-trained SAM2 model and adaptive test-time refinement (TTR). Specifically, a SAMSeg pre-training stage with a Multi-scale Feature Fusion Module (MFFM) is introduced to extract multi-scale semantic features from infrared images for visible-image generation. A denoising U-Net equipped with ResNet Attention (ResAtt) blocks is then employed to synthesize high-quality visible images. Furthermore, a Boundary- and Appearance-aware Test-Time Refinement (BATTR) strategy is developed to enhance robustness under limited-data conditions by jointly enforcing boundary preservation, semantic alignment, and visible-domain appearance constraints during inference without requiring paired visible references. Experiments on the DroneVehicle dataset demonstrate that SAMDiff-TTR achieves competitive image quality and consistently improves downstream object detection under small-sample settings, providing an effective cross-spectral data augmentation framework for infrared electromagnetic imaging and intelligent visual perception.