Adaptive Spatio-Temporal Vision Transformer for Online Action Recognition
Triet Nm. Kieu, Tan Hoa Nguyen, Viet Cuong Pham · 2024
Online Action Detection is a critical challenge in Computer Vision and Video Understanding, requiring the accurate identification of actions in streaming videos without access to future frames. Traditional methods, which heavily rely on 3D Convolutional Neural Networks for extracting spatiotemporal features, face limitations due to their complexity and high computational demands. Some methods treat videos as sequences of images and apply image encoders to locate and recognize actions, which becomes increasingly complex as the number of objects rises. Addressing these challenges, we propose a method that uses an image detection model to locate objects within frames. Unlike other methods that typically resize inputs to a fixed size, which can degrade performance, our approach handles flexible input aspect ratios. We then pack the input into sequences similar to natural language processing models and extract spatiotemporal features using a transformer model. By incorporating techniques from language models during training, our method achieves better performance compared to conventional Vision Transformer or CNN-based models.