An Improved MoViNet Algorithm for Lightweight Video Recognition
Zhou Rui-rui, Chen Xin · 2023
The MoViNets family offers accurate and efficient video recoginition networks deployed on mobile deveices, by alleviating the high memory consumption of 3D convolutional neural networks. The application of Stream Buffer greatly reduces the memory footprint, making MoViNet a lightweight video recognition framework for online inference and mobile devices. While causal operations including CausalSE are utilized to improve the MoViNet’s performance, however, the location information in spatial dimension is ignored. By incorporating the spatial dimension attention into the MoViNet, a lightweight Causal Spatio-temporal Attention(CSTA) is proposed to enhance video recognition efficiency. CSTA can not only embrace the causal relationship in time dimension to fit the Stream Buffer, but also pay attention to the spatial location information. The video classification experiment results show that the proposed lightweight attention module CSTA, increases MoViNet’s efficiency on the HMDB51 and UCF101 datasets, with a small computation budget.