Prediction Analysis of Greeting Gestures based on Data Augmentation and LSTM

Angga Wahyu Wibow, Eri Sato-Shimokawara, Kurnianingsih Kurnianingsih · 2024

The field of Human Activity Recognition (HAR) is growing significantly in several areas but little research focuses on cultural behavior. How machine learning can explain human activity as a promotional tool in understanding cultural differences in a region is very challenging. Studies using Recurrent Neural Network (RNN), especially Long Short-Term Memory (LSTM), require large data when used to research human activity recognition. Using small data is a challenge when using LSTM. This study aims to determine which data augmentation is most appropriate when used using LSTM to predict Japanese greeting gestures, namely eshaku, keirei, saikeirei, waving hands, and Indonesian greeting gestures, namely the movement of bringing both hands together in front of the chest with a slight bow of the head. This study proposes to compare several data augmentations, namely jittering, scaling, permutation, cropping, reverse, time shifting, frequency domain, and synthetic data generation. We evaluate model performance using the Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and coefficient of determination (R2) metrics. The experimental results show that Time Shifting has the best MSE, RMSE, and MAE values compared to others and Frequency Domain has the best R2values compared to others. This study concludes that Time Shifting and Frequency Domain are the best data augmentation methods when using LSTM in predicting greeting gestures using small data. These findings provide valuable insights into the effectiveness of various data augmentation methods in motion prediction tasks with limited data sets, as well as the importance of selecting appropriate data augmentation methods in human motion analysis. Future work could use data from two or more people using model explanations generated by machine learning and developing a new method for cultural behavior.

Read the paper · More papers on PaperTik