Video Action Recognition Model Based on Local and Global Spatial-Temporal Feature

Rongsen Wu, Jie Xu, Yuhang Zhang, Zixuan Li, Yiyao Li, Shixue Cheng · 2024

With the development of network technology, video applications have surged, necessitating timely review and supervision of the vast amount of uploaded data. Action recognition in videos, a key AI research area, has advanced through deep learning algorithms, utilizing both spatial and temporal features. Traditional methods like optical flow or 3D convolution are complex and computationally expensive. Difference modeling, however, improves model speed and accuracy. This paper designs video action recognition model based on spatial-temporal features, proposes a new difference information extraction method to extract local spatial-temporal features on each segment of the video. Then in the global stage, action rhythm feature is extracted with local feature difference. The final model achieved an accuracy of 97.6% on the UCF-101 dataset.

Read the paper · More papers on PaperTik