VIDA: Unsupervised Visible-to-Infrared Domain Adaptation for Object Detection Using Large Vision Language Model

Chanyeong Park, Junbo Jang, Jiyoon Lee, Jaehong Yoon, Minju Baek, Joonki Paik · 2025

In autonomous driving, reliable object detection across varied environmental conditions is essential, particularly when transitioning between spectral domains like visible light (RGB) and infrared (IR). This paper introduces a novel, fully unsupervised approach for RGB-to-IR domain adaptive object detection. By utilizing generative models to synthesize IR images from RGB inputs, our method eliminates the need for direct IR data collection. The proposed method falls under unsupervised domain adaptation. It generates IR images from RGB inputs to train a domain adaptive object detection model, leveraging a Large Vision-Language Model (LVLM) to transfer the target domain style using text prompts. By generating IR images solely from RGB inputs, the approach eliminates the need for expensive and time-consuming IR data collection, making it highly efficient. This enhances the robustness of object detection across spectral domains, improving vehicle safety in challenging environments. Extensive experiments demonstrate significant improvements in IR detection accuracy, all achieved without direct IR data during training, offering a cost-effective and scalable solution.

Read the paper · More papers on PaperTik