An Improved MoViNet Algorithm for Lightweight Video Recognition

Zhou Rui-rui, Chen Xin · 2023

The MoViNets family offers accurate and efficient video recoginition networks deployed on mobile deveices, by alleviating the high memory consumption of 3D convolutional neural networks. The application of Stream Buffer greatly reduces the memory footprint, making MoViNet a lightweight video recognition framework for online inference and mobile devices. While causal operations including CausalSE are utilized to improve the MoViNet’s performance, however, the location information in spatial dimension is ignored. By incorporating the spatial dimension attention into the MoViNet, a lightweight Causal Spatio-temporal Attention(CSTA) is proposed to enhance video recognition efficiency. CSTA can not only embrace the causal relationship in time dimension to fit the Stream Buffer, but also pay attention to the spatial location information. The video classification experiment results show that the proposed lightweight attention module CSTA, increases MoViNet’s efficiency on the HMDB51 and UCF101 datasets, with a small computation budget.

Read the paper · More papers on PaperTik