Cross-View Fused Network for Multi-View RGB-Based Action Recognition

Yingyuan Yang, Guoyuan Liang, Xinyu Wu, Biao Liu, Chao Wang, Jun Liang, Jianquan Sun · 2024

As the increasing number of multi-camera surveillance systems, researchers have been investigating the potential of multi-view action recognition. The skeleton-based methods are favorable due to the good performance. However, relying on specialized hardware with depth sensors or the efficient estimation for human poses, makes them less suitable for real-world deployment. In this study, we concentrate on the RGB modality for multi-view action recognition and propose a novel cross-view fused network. In details, gated adaptive fusion is utilized to aggregate the global representation across different views and spatio-temporal cross-attention module is employed to facilitate the effective fusion for spatio-temporal features. Based on three well-known multi-view action datasets: NTU-60, NTU-120, and N-UCLA, a series of experiments are conducted and verifies the effectiveness and robustness of the method under multi-view scenarios.

Read the paper · More papers on PaperTik