Light-aware luminance adaptive enhancement network for RGBT video object detection
Sugang Ma, Yafei Jiang, Liuxuanqi Gao, Yipei Li, Baojing Han · 2025
Visible-thermal Video Object Detection (RGBT VOD) aims to detect objects in RGB and thermal video sequences and predict their categories and locations. One of the core aspects of this task is how to efficiently and effectively perform the fusion of multiple modalities. Due to the limitations of the imaging mechanism, under poor lighting conditions, the image quality of the RGB modality will dramatically decline, reducing the detection performance. Supplementing with thermal data can effectively mitigate these challenges. To fully utilize the complementary information of the two modalities, existing works on feature fusion of RGB and thermal images require an extremely complex modality interaction process, which is time and resource-intensive. When converting the data of the two modalities to the YCbCr color space represented by luminance and color, only luminance information exists in the thermal image. In contrast, the RGB image is rich in luminance and color information. However, due to the imaging mechanism, the luminance information of the RGB modality can easily be affected by bad lighting. To simplify the design of the multimodal interaction module, reduce model complexity, and take full advantage of complementary multimodal information, we consider multi-modality fusion in terms of luminance information for multimodal data and propose a Light-aware Luminance Adaptive Enhancement Network (LA2ENet). Specifically, we design a Light-aware Luminance Adaptive Enhancement Module (LA2EM), which can sense the light information in the scene. When the luminance information in the RGB image is drastically affected due to bad lighting, the module will adaptively introduce the luminance information in the stabilized thermal infrared image to supplement and improve the quality of the RGB image. After that, we only use the luminance-enhanced RGB image as the input to the model, making full use of the modal complementary information while reducing the complexity of the model. We conduct extensive experiments on the VT-VOD50 dataset, compared to the baseline, our LA2ENet improves the AP50 metric by 4.66% and is almost equal to the baseline in detection speed, which demonstrates the effectiveness and efficiency of our proposed method.