Can We Trust Large Language Models for Video Analysis: An Exploration of Hallucination in Multimodal LLMs

Yiqiu Zhou, Jina Kang, Ha Nguyen · Proceedings. · 2025

Multimodal Large Language Models (MLLMs) show promise for facilitating qualitative video analysis, potentially reducing the manual effort required for video selection and coding.However, MLLMs are subject to hallucination, where models generate incorrect or unfounded descriptions and interpretations.Current frameworks for evaluating MLLM hallucinations primarily focus on object-centric tasks, with limited focus on behaviors and social interactions common in education data.Applying MLLMs to analyzing 36 video clips of collaborative learning activities, we identify three distinct types of hallucinations: description hallucination (factual inconsistency), interpretation hallucination (misinterpretation of behaviors), and instruction hallucination (deviation from analytical focus).Despite MLLMs' ability to process educational videos, they exhibit gender misidentification and inaccurate interpretation of student roles and behaviors.This exploratory work develops a domain-specific framework to understand and evaluate MLLM hallucinations, laying the groundwork for more reliable implementation in educational research and calling for further investigations.

Read the paper · More papers on PaperTik