DenseNet-Transformer Model: A Hybrid Architecture for Improved Human Action Recognition in Videos

R. Athilakshmi, Lakshmi Gnaneswara Prasadu Indugapalli, Anusuri Bhuvan Sai, M. Chandra Kiran Teja · 2025

Human Action Recognition (HAR) in video sequences is a critical yet challenging task in computer vision, requiring the capture of both spatial and temporal features for accurate classification of videos. While traditional Convolutional Neural Networks (ConvNets) have achieved considerable success in image recognition, they often struggle with HAR due to their focus on spatial features and inability to effectively model temporal dynamics. To overcome this limitation, we propose a novel DenseNet-Transformer model that combines the strengths of both DenseNet and Transformers. DenseNet is used for robust feature extraction from video frames, while Transformers are leveraged for their powerful sequence modeling capabilities to capture temporal dependencies across frames. Our hybrid DenseNet-Transformer architecture achieves state-of-the-art performance on benchmark datasets, with 93.75% accuracy on the UCF-101 dataset and 87.5% accuracy on the Kinetics dataset. These results highlight the effectiveness of our approach in enhancing the accuracy and reliability of HAR systems.

Read the paper · More papers on PaperTik