Two stage Multi-Modal Modeling for Video Interaction Analysis in Deep Video Understanding Challenge
Siyang Sun, Xiong Xiong, Yun Ping Zheng · Proceedings of the 30th ACM International Conference on Multimedia · 2022
Interaction understanding between different entities in human-centered movie video is receiving more and more attention. Recently, a deep video understanding (DVU) task is proposed to identify interactions between different person entities on scene level task. However, limited samples of DVU dataset and multiple complex interactions make it difficult. To tackle these problems, we propose a two stage multi-modal method to predict the interaction between entities. Specifically, scene segment is first divided into several sub-scene clips, meanwhile face trajectory and person trajectory are obtained through face tracking/recognition and skeleton-based person tracing. Then, we extract and jointly train multi-modal features in the same semantic space including face emotion feature, person entity feature, visual features, text features and audio features. Finally, zero-shot transfer model and multiple classification model are proposed to predict interactions together. The experimental results show that our method performs new state-of-the-art on the DVU dataset.