VARSew Benchmarking: A Visual Action Recognition on Garment Sewing using Various Deep learning models
Aly E. Fathy, Ahmed Hamza Asad, Ammar Mohammed · 2023
The method of analyzing human action using computer and machine vision technologies is known as human action recognition (HAR). In recent years, research in this field has made substantial strides in a variety of application domains. However, HAR is still constrained in a number of fields because of the lack of sufficient, sizable training datasets. Many factors impact the availability of suitable HAR data in specific domains. Garment sewing is one of those interesting industrial fields. To demonstrate the growth of the global garment sewing business, several international reports were made public. As far as we know, HAR data for the garment sewing industry is limited. Thus, in this work, we benchmark and validate VARSew, a new dataset covering human garment sewing actions. VARSew is designed to capture multiple research challenges in HAR. We benchmark VARSew on two types of video classification tasks: binary and multi-class. VARSew contains 3,121 video segments of 49,936 frames. We conducted extensive benchmark experiments with 32 state-of-the-art HAR models and compare their performance. The results on both binary and multi-class reveal that the best accuracy scores are 80% and 95% for binary and multi-class classification, respectively. As a result, deep learning vision transformer (Vit) models outperformed deep learning convolutional and ResNet models on both the VARSew multi-class and VARSew binary-class datasets.