Deep Learning for Folk Dance Classification: An Analysis of ConvLSTM and LRCN Architectures
Joel Thaduri, S Chakradhar Amingad, Ganesh Bhaiyya Regulwar, Vedha Thirmanpally, Jaiwanth Reddy · 2024
This study looks into the classification of various folk dances from video sequences using methods for deep learning. Knowing that Long Short-Term Memory Recurrent Convolutional Network (LRCN) and Convolutional LSTM (ConvLSTM) designs are so good at capturing temporal and spatial relationships in video data, the study focuses on using them. A diverse dataset of folk-dance videos from different geographic regions and cultural contexts was curated and preprocessed into sequential image frames for model training. The ConvLSTM and LRCN models were constructed and trained using the TensorFlow/Keras framework on GPU-accelerated hardware. In this paper, analysis revealed effective learning during training, with both models demonstrating increased training accuracy and decreased training loss. However, significant fluctuations in validation accuracy and loss indicated potential overfitting and instability in generalization to unknown data. To deal with these problems, we suggest employing regularization strategies like dropout, early stopping, and further hyperparameter tuning. Future work could explore transfer learning and more diverse validation procedures to enhance model performance. Despite the challenges, this study provides a foundational approach for improving video classification systems in the rich and varied domain of folk dances.