ViTHAR: Vision Transformers for Human Action Recognition in Videos

Divya Rani R, C. J. Prabhakar · 2024

Human action recognition is a popular research topic that has drawn the interest of numerous researchers due to its potential applications in the fields of video surveillance, healthcare, human-computer interaction, and sports analysis. But there are a lot of challenges associated with the action recognition, some of these difficulties include differentiation of similar actions, solving occlusion issues, background clutter, and adapting to changes in viewing angle. Recently, Vision Transformers (ViT) have demonstrated exceptional performance in action recognition, which represents long video sequences captured across spatial and temporal indexes in a video. In this paper, a method for human action recognition using Vision Transformer models referred as ViTHAR is proposed to enhance the accuracy and efficiency of action recognition in video sequences. Unlike traditional convolutional neural networks, ViTs take advantage of self-attention mechanisms to better capture spatial and temporal relationships. To demonstrate the efficacy of the proposed method, we conducted experimentation on two benchmark datasets, UCF-101 and HMDB-51 datasets, and compared the proposed ViTHAR method with various state-of-the-art HAR methods. The experimental findings reveal that the proposed ViTHAR method performed exceptionally well compared to other state-of-the-art HAR methods. The proposed ViTHAR method achieves the highest recognition accuracy of 96.53% and 79.45% for UCF-101 and HMDB-51 datasets respectively, using the Vision Transformer model variant ViT-Base. The experimental findings acquired from the proposed ViTHAR method show that it performed effectively in dealing with all of the challenges that are commonly encountered in action videos.

Read the paper · More papers on PaperTik