Action Recognition Using Dynamic Mode Decomposition for Temporal Representation

Kyle Pawlowski, Sumit Chakravarty, Ying Xie, Arjun Kumar Joginipelly · 2020

This paper explores a new method for representing temporal information found in videos. Dynamic Mode Decomposition (DMD), a method commonly used to reduce the computational effort for other big-data tasks such as flow calculations, is used in this study to aggregate changes between multiple frames. This is applied to the challenge of human action recognition (HAR) tasks using a Two-Stream architecture. Such an architecture takes two convolutional neural networks (CNNs), one analyzing the spatial data and the other analyzing the temporal data as calculated using DMD. This method is compared against others using two common benchmarks, the UCF-101 dataset and the HMDB-51 dataset, achieving 46.0% accuracy on the UCF-101 and 30.8% on the HMDB-51.

Read the paper · More papers on PaperTik