MI-Cap: A Multi-Modal Interpretable Model for Video Captioning
Antoine Hanna-Asaad, Decky Aspandi-Latif, Titus B. Zaharia · 2025
Video captioning aims to describing video content in natural language, which implies understanding and interpreting scenes, objects, actions and events that are present in the video. Current approaches have mostly concentrated on visual cues, often neglecting the rich information available from other important modalities, such as the audio or textual (e.g. subtitles) channels. In this paper, we introduce a novel video captioning method trained with a multi-modal contrastive loss function that emphasizes both multi-modal integration and interpretability. Our approach is designed to capture the rich dependencies between the different modalities, resulting in more accurate and pertinent captions. Concerning the interpretability issues, by exploiting multiple attention mechanisms, the model is able to provide explanations of the results proposed. The experimental evaluation, caried out on widely used benchmark datasets such as MSR-VTT and VATEX, demonstrate that the proposed method performs favorably against state-of the-art models.