Evaluating Data Quality and Preprocessing Methods to Enhance Skeleton-Based Action Recognition in Retail Environments
Samer Yousef, Chengjun Han, Naoya Chiba, Koichi Hashimoto · 2025
The accuracy of skeleton-based action recognition models can be significantly improved using data processing techniques, particularly in complicated environments such as retail stores where people’s activities are dynamic and prone to noise and occlusions. This paper investigates how various preprocessing techniques, including smoothing, data compensation, and noise augmentation, influence the accuracy of two action recognition models, STGCN and STGCN++. We evaluated these methods using different data sources: key points of YOLO, HRNet, and motion capture (MoCap). Our findings demonstrate that higher quality data, such as projected 2D MoCap, lead to significantly better model performance, with STGCN++ achieving a mean accuracy of $91.26 \%$, compared to $81.06 \%$ with YOLO key points. We also show that applying preprocessing methods to noisier data, such as YOLO, can reduce performance degradation, improving their robustness for real-world applications. These results highlight the importance of preprocessing in implementing robust action recognition systems that can operate in dynamic, real-world settings and open the way for further improvements in the field.