STGA-Net: Spatial-Temporal Graph Attention Network for Skeleton-Based Temporal Action Segmentation
Xiaoyan Tian, Ye Jin, Zhao Zhang, Peng Liu, Xianglong Tang · 2023
Temporal action segmentation aims at dense labeling of video frames with a series of action classes in long and untrimmed videos. However, previous methods heavily rely on generating an initial prediction with temporal convolutional layers and refining the predictions over the following stages based on RGB features. This results in the lack of an explicit action segment transition rule and loss of high-level semantic information, such as complex spatial-temporal correlation among the human joints between frames. Therefore, we present a spatial-temporal graph attention network (STGA-Net) for skeleton-based temporal action segmentation. In particular, we propose a spatial-temporal attentive block for prediction generation, which adapts an encoder-decoder architecture, where both the encoder and decoder contain various graph spatial-temporal attention blocks to model the dynamic and non-linear correlation among joints. Experiments on three challenging datasets (PKU-MMD, HuGaDB, and LARa) demonstrate that the performance of our STGA-Net exceeds that of the state-of-the-art and alleviates over-segmentation and ambiguous boundary errors to a large degree.