Comparative Analysis of Human Activity Recognition Using Deep Neural Networks

Syed Ali Faraz Kazmi, Athar Ali, Faiza Hussain, Halika Hameed, Syed Ashar Ali · 2024

Human Activity Recognition (HAR) has emerged as a vital field within computer vision, aiming to develop intelligent systems capable of automatically identifying and classifying human actions [1]. Although there are advancements in Deep learning techniques but there is gap in understanding the comparative effectiveness of Deep Learning Models such as, Convolutional Neural Networks (CNNs) and Vision Transformers (ViT) [2]. To cover this gap this article delves into the application of deep learning techniques, CNNs and ViT models, for achieving accurate and robust HAR. This article presents the comparative analysis of CNNs and ViT. CNNs excel at capturing spatial features within individual video frames, enabling them to recognize subtle nuances of human movement [3]. ViT models, on the other hand, excel at analyzing global relationships within the frames, providing a broader understanding of the overall action [3]. The chosen deep learning models are employed to analyze video frames and extract relevant features, ultimately leading to accurate classification of various human activities. This research contributes to the field of HAR by comparing the effectiveness of CNNs and VIT models for enhanced recognition performance. Results depicts that the CNN model achieved an accuracy of 94 %, while the VIT model achieved an impressive 95 % accuracy when trained on sports dataset.

Read the paper · More papers on PaperTik