Multi-Modal Human Action Segmentation Using Skeletal Video Ensembles
James Dickens, Pierre Payeur · 2023
Beyond traditional surveillance applications, sensor-based human action recognition and segmentation responds to a growing demand in the health and safety sector. Recently, skeletal action recognition has largely been dominated by spatio-temporal graph convolutional neural networks (ST-GCN), while video-based action segmentation research successfully employs 3D convolutional neural networks (3D-CNNs), as well as vision transformers. In this paper, we argue that these two inputs are complementary, and we develop an approach that achieves superior performance with a multi-modal ensemble. A multi-task GCN is developed that can predict both frame-wise actions as well as action timestamps, allowing for the use of fine-tuned video classification models to classify action segments and achieve refined predictions. Symmetrically, a multi-task video approach is presented that uses a video action segmentation model to predict framewise labels and timestamps, augmented with a skeletal action classification model. Finally, an ensemble of segmentation methods for each modality (skeletal, RGB, depth, and infrared) is formulated. Experimental results yield 87% accuracy on the PKU-MMD v2 dataset, delivering state-of-the-art performance.