Deep Learning Techniques and Behavior Recognition in Video Image Sequence Analysis

Dan Zeng · 2024

Aiming at the performance bottleneck of distinguishing complex scenes and similar actions in video behavior recognition, a hybrid neural network architecture based on spatio-temporal feature fusion is constructed in this study. By organically integrating the spatial perception advantages of 3D convolutional neural network and the time series modeling capabilities of long-term and short-term memory networks, a double-stream feature interaction framework is innovatively designed: 3D-CNN is used to analyze the spatial topological relationship of objects in video frames, and LSTM is used to capture the time series law of cross-frame motion evolution. In the feature fusion stage, an attention weighting mechanism is introduced to dynamically adjust the contribution weight of spatio-temporal features, forming a multi-level behavior representation. After datasets testing, the average recognition accuracy of the model on UCF101 and HMDB51 benchmarks reaches 94.7% and 72.3%, respectively, which is 5.2% higher than that of the benchmark method. The validation experiment of the module shows that the combined training strategy of spatio-temporal features can reduce the recognition error rate of the model for occluded scenes by 31%, while applying the attention mechanism can improve the article discrimination of similar actions by 19.8%. This study provides a more interpretable feature fusion paradigm for dynamic visual understanding.

Read the paper · More papers on PaperTik