Scene-aware Graph-enhanced Multimodal Collaboration for Video-grounded Dialogue
Shan-Shan Du, Hanli Wang · 2025
Video-grounded dialogue systems face inherent chal- lenges in explicitly modeling spatiotemporal reasoning and handling complex scenarios involving multi-event interactions and long-term dependencies. To address these limitations, a novel scene-aware graph-enhanced multimodal collaborative model (SGMCM) for video-grounded dialogue is proposed, which collaborates explicit structural reasoning with implicit sequence modeling. The framework comprises three core components: (1) a fine-grained explicit feature extractor encodes entity relationships in video frames and dialogue sequences via separate visual and textual scene graphs, and refines intra-structure features via graph convolutional networks; (2) an implicitly aligned feature extractor, built upon adapted BART model, encodes video, audio, and text sequences into aligned latent representations; and (3) an explicit-implicit feature fusion module is designed, which adaptively integrates graph-structured explicit features with pre- trained BART implicit features via a dynamic gated fusion method, followed by cross-modal joint representation learning using multi-head attention. Experiments on the public datasets of AVSD-DSTC7, AVSD-DSTC8, and NExT-QA demonstrate the superiority of our SGMCM over existing methods, with ablation studies confirming each component’s contribution.