Graph Convolutional Networks With Minimal Appearance Information For Action Recognition
Hiroaki Tani · 2024
Skeleton-based action recognition, specifically methods using Graph Convolutional Networks (GCN), have gained attention in recent years due to its low computational complexity. However, skeleton-based action recognition, dealing only with joint coordinates, is unable to recognize actions that involve the use of objects. On the other hand, appearance-based action recognition, specifically methods using 3D Convolutional Neural Networks (CCN) can utilize object information, but it is prone to high computational complexity and domain gaps. To address these, we propose a method which efficiently combine skeleton-based and appearance-based action recognition. By utilizing the temporal attention obtained from a GCN model, we extract a single image from a video and incorporate the necessary bare minimal of appearance information. Our method can be applied to many existing GCN models, and it is capable of enhancing their performance. Using datasets NTU-RGB+D, NTU-RGB+D 120, and IKEA-ASM, we confirm the effectiveness of our method with accuracy improvements of up to 8.13 percentage points.