MotionMLP: End-to-End Action Recognition with Motion Aware Vision MLP

Xiangning Ruan, Zhicheng Zhao, Fei Su · 2023

Action recognition aims to interpret complex spatiotemporal patterns in the video. Current methods utilize CNN or Transformer structures, requiring extensive pre-training methods and optical flow to capture motion information. Such approaches are computationally expensive, necessitate significant storage, cannot be trained end-to-end, and typically neglect joint learning of temporal and spatial streams. In this paper, we propose MotionMLP, a novel MLP architecture that extracts motion information from videos, then dynamically adjusts the connection between tokens and static weights within the MLP structure. The MotionMLP solely relies on video frames as input and is independent of any pre-training method or optical flow computation. The experimental results indicate that MotionMLP outperforms the previous SOTA real-time end-to-end methods on UCF101 and HMDB51, and relies on one-tenth of the parameters compared with typical two-stream CNN approaches, while operating ten times faster.

Read the paper · More papers on PaperTik