Enhancing Human Action Recognition in Videos through Dense-Level Features Extraction and Optimized Long Short-Term Memory
Najmul Hassan, Abu Saleh Musa Miah, Jungpil Shin · 2024
Human action recognition (HAR) in videos is a field currently receiving considerable attention in computer vision and pattern recognition. While numerous researchers have been working on developing a HAR system, they still face challenges in achieving satisfactory performance in video-based activity recognition systems. Video-based action recognition poses a significant research challenge due to temporal feature dependency. Currently, there is a demand to develop dynamic HAR systems that can deliver high-performance accuracy with generalizability. In order to address these shortcomings, we propose a dynamic HAR system by leveraging a deep learning (DL) based temporal feature extraction approach. In this process, we initially use the DL-based DenseNet121 model to extract frame-level dense features. Subsequently, we feed these features into an optimized Long Short-Term Memory (LSTM) network to learn dependencies and process data for optimal predictions. Furthermore, during the testing phase, an iterative fine-tuning procedure is incorporated to update the high parameters of the trained system, aiming to achieve optimal results. Experiments conducted on a benchmark dataset demonstrate superior performance in terms of both accuracy and losses compared to existing methods.