Hallucination Mitigation in LLM-Based Video Captioning via Multi-Scale Feature Fusion

Jiancheng Zhou, Ye Wang, Qun Liu · 2024

With the development of visual representation methods and pre-trained language models, video captioning technology has made significant progress. However, when describing objects and actions in videos, models are still prone to generating hallucinations, which introduce information that is irrelevant to the actual video content. These hallucinations greatly limit the practical applications of video captioning.A major issue is that the visual encoders widely adopted in vision-language models are derived from CLIP. Although CLIP has demonstrated outstanding performance in various visual understanding tasks, it still faces limitations in effectively capturing comprehensive visual information, such as fine-grained visual semantics. Additionally, given that CLIP is trained using pairs of images and text, directly using CLIP as a video encoding module can also lead to the model’s inability to effectively capture the spatial-temporal information of videos. To overcome these difficulties, we introduce a multi-scale feature modeling network that utilizes visual feature fusion for video captioning. First, we compensate for the lack of fine-grained visual semantics caused by relying solely on a single pre-trained model by incorporating additional salient region features. Next, we conduct both local and global spatialtemporal modeling across consecutive frames to capture global dependencies and local action information within the video. Finally, a multi-scale feature fusion approach is employed to achieve unified semantic modeling of the video, facilitating the sentence decoding process in LLM. We conducted both quantitative and qualitative evaluations of our approach, demonstrating the significant potential of this novel method. Compared with the baseline method, the text description we generates more effective at mitigating hallucinations related to objects and action descriptions.

Read the paper · More papers on PaperTik