Object Action Recognition Algorithm Based on Asymmetric Fast and Slow Channel Feature Extraction
Lingchong Fan, Yude D. Wang, Zhang Ya · 2024
Aiming at the problems of insufficient video motion feature extraction and low recognition accuracy of the existing symmetric object action recognition model, we propose an asymmetric video object action recognition model that extracts content features from a single slow channel and motion features from two fast channels. A single slow channel extracts video content information; two fast channels extract the motion features of the video and uses the lateral connection to fuse the features with the slow channel. We obtain the object action recognition features after the extracted content features and motion features through global average pooling and concatenation operation, so as to realize the action recognition of motion objects. We conduct comparative experimental studies of different model architectures on the UCF101 and HMDB51 datasets. The experimental results show that the recognition results of the asymmetric object action recognition models Top1 and Top5 on UCF101 dataset are 1.32% and 1.51% higher than those of the symmetric SlowFast model, 4.44% and 4.05% higher on HMDB51 dataset, respectively. In addition, the influence of fast channels with different frame rates on the recognition results is studied in the constructed model, and the experimental results show that the recognition effect is the best when the frame rate α1is equal to 4 and α2is equal to 8. Our proposed asymmetric model achieves good performance in object action recognition, which makes a clear contribution to these improvements.