The Video Behavior Recognition Based on R (2+1) D
Xingjie Xu, Weihua He · 2023
In the behavior recognition task, sometimes only a single frame image can be used to accurately identify behavior, but in order to achieve high overall accuracy, time-scale feature fusion is essential, some articles propose to use R(2+1)D network structure, the feature extraction of single-frame image and time-scale feature fusion into independent two steps, this method not only reduces the scale of the model, but also improves the recognition accuracy. This paper reproduces and analyzes the representative article “A Closer Look at Spatiotemporal Convolutions for Action Recognition”, and achieves a higher accuracy of 0.73 than the original 0.666 on the HMDB51 dataset.