LLM based semantic fusion of infrared and visible image caption

Kai Wang, Shuli Lou, Jinhong Wang, Xiaohu Yuan · 2024

Conventional image captioning methods are mostly based on visible image, which perform poorly in low-light scenarios. To address this issue, we propose a semantic fusion approach for infrared and visible image captioning based on large language model (LLM). Firstly, we construct a dataset containing both infrared and visible images to train the caption model and obtain image captions corresponding to both infrared and visible modalities. Secondly, we design an enhanced caption module to incorporate the results of object detection as supplementary information, combining them with image captions using LLM to improve the accuracy of the model’s image captions. Finally, to fully leverage the complementary nature of the infrared and visible light modalities, we design a multimodal fusion module to semantically fuse the image caption results from both modalities using LLM. Experimental results demonstrate that our method outperforms the baseline on multiple metrics and maintains good performance in complex scenarios.

Read the paper · More papers on PaperTik