Multimodal-Semantic Context-Aware Graph Neural Network for Group Activity Recognition
Tianshan Liu, Rui Zhao, Kin‐Man Lam · 2021
Group activities in videos involve visual interaction contexts in multiple modalities between actors, and co-occurrence between individual action labels. However, most of the current group activity recognition methods either model actor-actor relations based on the single RGB modality, or ignore exploiting the label relationships. To capture these rich visual and semantic contexts, we propose a multimodal-semantic context-aware graph neural network (MSCA-GNN). Specifically, we first build two visual sub-graphs based on the appearance cues and motion patterns extracted from RGB and optical-flow modalities, respectively. Then, two attention-based aggregators are proposed to refine each node, by gathering representations from other nodes and heterogeneous modalities. In addition, a semantic graph is constructed based on linguistic embeddings to model label relationships. We employ a bi-directional mapping learning strategy to further integrate the information from both multimodal visual and semantic graphs. Experimental results on two group activity benchmarks show the effectiveness of the proposed method.