Behavioral Recognition Based on Representation Learning in Video Scenes
Dongjun Zhu, Bidare Divakarachari Parameshachari · 2022 IEEE 2nd International Conference on Mobile Networks and Wireless Communications (ICMNWC) · 2022
With the construction of smart cities in China, surveillance cameras have become ubiquitous, which has also caused an unprecedented expansion of surveillance video data. Using the intelligent video surveillance system developed by behavior recognition technology to replace human beings for video surveillance, which can reduce human resources and time resources, and has great practical value. This paper mainly explores the algorithm principle of the video behavior recognition method based on representation learning. By constructing a dual-stream network, we extract the appearance information and action information in the video respectively, and then fuse the two kinds of information to classify the behavior on the fusion results. Compared with the single-stream, the dual-stream network broadens the dimension of network learning and partly enriches the capture of information in video sequences, but it lacks the extraction of long-distance information in video sequences. Therefore, a two-stream network structure based on temporal fragment fusion is improved based on the ordinary two-stream network. The network first divides the overall behavior into multiple action segments according to the optical flow intensity, and then classifies the temporal features and modal features of the different segments after fusion.