Swin Transfomer with structure from motion for group activity recognition

Jiajun Huang, Ling Chen, Yunhao Yao, Jiawei Xu, Chuanxu Wang · Journal of Electronic Imaging · 2025

We aim to identify individual actions and group activities from videos. Existing solutions primarily model spatial and temporal relationships based on the positional information of individual participants, but there is a clear limitation: due to the relative motion of the camera and the discrete nature of pixels, the pixel coordinate sequences of characters in videos often fail to accurately reflect their true location information in the real physical world, and no research team has yet proposed an effective solution to this problem. To overcome this challenge, we propose a method based on structure from motion (SFM). We utilize SFM, combined with deep learning which is called sub-bundle adjustment for secondary correction of three-dimensional pose keypoint information as part of the input of the model, and employ transformer and Swin Transformer architectures to richly represent local and global spatial features; then, we designed a multi-channel feature fusion module. This module integrates local color features (RGB), global RGB features, and pose features through an adaptive weighting and fusion approach, thereby providing a more comprehensive feature representation for subsequent group behavior recognition. Finally, followed by long short-term memory to capture temporal feature dependencies at the video level. In our empirical study, we explored multiple approaches to fuse these feature representations and confirmed the complementary advantages among them. On the Volleyball dataset, our method achieves 93.9% under multi-class accuracy (MAC) indicator and 94.3% under mean per class accuracy (MPAC). On the Collective dataset, our method achieves 95.8% under MAC metric and 95.9% under MPAC. Experimental results demonstrate the importance of motion recovery and deep learning–based secondary reconstruction result correction. More importantly, our proposed method has achieved significant superior performance on recognized benchmarks for group activity recognition.

Read the paper · More papers on PaperTik