Motion Embedded Image: A combination of spatial and temporal features for action recognition

Quang-Tri Le, Nham Huynh-Duc, Chung Thai Nguyen, Minh–Triet Tran · 2022

Demand for human activity recognition from videos has rapidly increased in many real-life applications, e.g., video surveillance, entertainment, healthcare, child and old homes, etc. Moreover, the explosion of short-form videos on social networking platforms such as Tiktok, Facebook, Youtube, etc., makes this problem gain much more attention. In this paper, we focus on the problem of human activity recognition in general short videos. Compared with still images, clips provide both spatial and temporal information, and the challenge is to capture the complementary information on appearance from still frames and motion between frames. Our contribution is two-fold. First, we study an approach of using motion embedded Image in a variation of two-stream ConvNet architecture: one stream is a motion stream capable of capturing and recognizing motion based on embedded batches of frames; another one is a normal image classification ConvNet being fed still frames to classify static appearance and recognize the missing spatial information from the first stream. Second, we build a brand new dataset of Southeast Asian Sports short videos, consisting of both standard videos with no effect and non-standard videos with effects - a modern factor that all currently available datasets being used for benchmarking models lack. Our model is trained and evaluated using different backbone architectures and on two benchmarks: UCF-101 and SEAGS-V1. The result shows that this is a model with competitive performance compared to previous attempts to use deep nets for human activity recognition in short-form videos.

Read the paper · More papers on PaperTik