A Few-Shot Learning for HAR Using 3D Siamese Network
Hua Guanghui, Govindaraju Hemantha Kumar, V. N. Manjunath Aradhya, Punithavalli · 2024
In artificial intelligence, deep learning techniques were applied in a wide field due to their ability to automatically learn complex patterns from raw data, such as image classification, object detection, image generation, and video analysis in computer vision. These applications are just the tip of a substantial iceberg. With the development of deep techniques to achieve state-of-art accuracy, learn precisely relevant features from raw data, and improve the robustness of models to make them versatile solutions. The traditional deep learning techniques require vast computational resources, including high-performance computer GPUs or TPUs, and can be expensive in terms of hardware. In Human Activity Recognition Systems (HARs). The video needs more computational resources than image processing. In this paper, we proposed a few-shot learning approach based on a 3D Siamese Network with depth video sequences to address the action video similarity or dissimilar problem. Convolutional Neural Networks (CNNs) are well-suited for images and spatial data processing. The 3DCNN is an extension of CNNs to incorporate an additional temporal dimension to capture the temporal pattern. From the 3D Siamese network, we extract more effective features with spatial-temporal action information from raw data. We employed the 3DCNNs Siamese Network (3DCNNs+SN) with shared weight to recognize or classify a new class of data without new sample data only based on learning training existing sample constructed model. In addition, we divided the benchmark dataset MSR Action 3D dataset into sub-dataset1 (SD-1), sub-dataset2 (SD-2) and sub-dataset3(SD-3) to train and validate our proposed method. For dissimilar sub-dataset 2 we got a 98.1% accuracy, 33% on high-similar sub-dataset 3 and $\mathbf{63.5\%}$ on merged sub-dataset SD-4 from sub-dataset2 (SD-2) and sub-dataset3(SD-3).