Combining Deep Learning Techniques for Enhanced Human Activity Recognition: A Hybrid CNN-LSTM Fusion Approach
M. Prabu, Rupesh Naidu M, P N Asif · 2024
Over recent years conventional pattern recognition pattern recognition techniques have advanced significantly. Nonetheless, these methods heavily depend on manual feature extraction, potentially limiting the overall performance and generalizability of the model. Given the rising prominence and efficiency of deep learning approaches, there has been a growing interest in leveraging these methods for human action recognition in mobile and wearable computing contexts. Despite this enthusiasm, existing deep learning architectures may not fully capitalize on the complex spatio-temprial patterns inherent in human actions. Moreover, deploying deep learning models in real-world scenarios often demands robustness, efficiency, and interpretability, which remain significant challenges. To address these issues, this paper proposes a novel approach: a deep neural network that integrates convolutional layers with long short-term memory(LSTM) units. By fusing the strengths of CNNs in spatial feature extraction with the temporal modelling capabilities of LSTMs, our hybrid architecture aims to enhance the accuracy and robustness of human activity recognition systems. This fusion enables the model to automatically learn hierarchical representations from data, eliminating the need for manual feature engineering. The approach holds, eliminating the need for manual feature engineering. The approach holds promise for applications in healthcare, sports analysis and assistive technology, where accurate and reliable HAR is essential. Overall, The proposed model achieves accuracy of 80% and 92% on the UCF50 dataset, a comprehensive video dataset, utilizing the CNN and LSTM models individually, respectively. These findings indicate that our hybrid model exhibits superior robustness and enhanced activity detection capabilities compared to certain previously reported results of only sensor data, surpassing an 90% accuracy threshold when trained on video-based data. Notably, the model demonstrates adaptability in extracting activity features whie maintaining a streamlined parameter count and achieving heightened.