Cross-modal fusion and temporal enhancement for egocentric action recognition
Dengdi Sun, Xueliang Zhang, Bin Luo, Zhuanlian Ding · 2024
Egocentric action recognition is an important research topic in the field of video understanding, which focuses on parsing and understanding the actions of characters in videos. In this paper, we propose a novel egocentric action recognition method, which deeply mines the text information within the labels and closely combines the video data with the text information, thus achieving a more comprehensive and detailed behavior pattern capture. By incorporating the dual-flow network architecture of RGB and Optical flow, our method can more accurately capture the temporal dynamic information in videos, thereby improving the spatio-temporal representation of our model in egocentric action recognition. In order to verify the effectiveness and generality of the proposed method, we performed experiments on several commonly used first-view datasets (such as EGTEA and GTEA gaze+). The results demonstrate that compared with the state-of-the-art, our method achieves significant improvements in both key metrics, accuracy and average class accuracy.