Visual Evidence-Aware for Object Hallucinations Rectification in LLM-Based Video Captioning
Ye Wang, JianCheng Zhou, Qun Liu, Feng Hu, Guoyin Wang · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Recent neural models for video captioning are typically built using a framework that combines a pre-trained visual encoder with a large language model(LLM) decoder. However, large language models in video captioning often generate non-existent entities, known as object hallucinations, which severely limit performance. To mitigate object hallucinations, two key issues remain: 1. Biased training data and Knowledge bias in LLM leads models to generate hallucinations; 2. Current methods focus on removal rather than restoring the correct visual content, reducing caption completeness. To address these issues, we propose a visual evidence-aware for object hallucination rectification in LLM-based video captioning. Generally, our model aims to diagnose and correct those generated object hallucinations, and then supplement missing visual content by constraining the process of text description generation. Specifically, we first generate captions by words based on the input video. When decoding each object description, the decoder utilizes visual features for hallucination diagnosis and correction, proposing visual evidence to modify hallucinatory descriptions. This process ensures the generated captions align with the visual content, alleviating the generation of object hallucinations. Compared with the baseline models, our method performs state-of-the-art performance in video captioning, especially avoiding neglecting objects in the visual content caused by the generated hallucinatory descriptions.