Leveraging Language-Aligned Visual Knowledge for Remote Sensing Image Spectral Super-Resolution
Bowen Chen, Keyan Chen, Liqin Liu, Zhenwei Shi, Zhengxia Zou · 2024
Spectral super-resolution (SSR) can reconstruct hyperspectral images (HSIs) from images with fewer spectral bands, such as RGB images, to furnish critical spectral information for var-ious downstream applications. However, existing SSR meth-ods for remote sensing images often focus solely on mapping RGB pixels to spectral signatures, neglecting semantic con-tent. This oversight may result in the omission of critical semantic details and inaccurate spatial-spectral textures. To address these issues, we propose Language-Aligned Visual Embedding-Driven Spectral Super-Resolution (LAVER), a novel SSR method that leverages the rich vision-language pri-ors in the Contrastive Language-Image Pre-training (CLIP) model. Specifically, we first design a denoising diffusion model-based baseline and then propose Local CLIP and Language-Enhanced Spectral Fusion Block (LESFB) to un-leash the potential of language-aligned visual embeddings. LAVER can reconstruct HSIs with higher fidelity and better details than the baseline and other advanced methods. Codes can be found at http://github.com/Mr-Bamboo/LAVER.