SlowFast Action Recognition Network Based on Improved Residual Structure

Guoqing Liu, Qian Wang, Yuan You, Hao Li · 2024

The SlowFast action recognition network, built upon the ResNet3D-50 backbone, excels at capturing video information across different temporal scales. However, it falls short in leveraging global spatiotemporal context information when handling complex spatiotemporal features, which hampers its ability to model fine-grained changes and long-term dependencies in actions. To address these issues, this study introduces the Contextual Transformer (CoT) module into the SlowFast network. By integrating static context encoding and dynamic context learning, the CoT module enables efficient attention to global spatiotemporal features. It replaces the 1×3×3 convolutions in some residual blocks of ResNet3D-50, resulting in the proposed CoT-SlowFast network. Experiments conducted on the Atomic Visual Actions(AVA) dataset focused on eight action categories (four person interaction actions and four object manipulation actions). The results demonstrate that the CoT-SlowFast network improves accuracy by 1.39% for person interaction actions and by 0.32% for object manipulation actions, with an overall average improvement of 0.88% compared to the SlowFast network. In summary, the introduction of the CoT module into the CoT-SlowFast network effectively enhances action recognition accuracy, validating its efficacy in modeling complex spatiotemporal features.

Read the paper · More papers on PaperTik