MMST-GCN+: A Multi-Modal ST-GCN+ for Computer-Aided Education by Introducing a Sensor Skeleton
Ziyun Zhao, Aohua Song, Qingyun Xiong, Siyu Zheng, Junqi Guo · 2023
Multi-model data can improve the performance of image-based models for computer-aided education (CAE) by utilizing features from more dimensions rather than images only. The spatial temporal graph convolution network (ST-GCN) is a two-stream model for classifying human actions based on their skeleton data. This work proposes a novel intelligent data model (MMST-GCN) based on ST-GCN for multi-model image-based CAE. MMST-GCN integrates image and sensor data into a unified framework by introducing a sensor skeleton that provides stronger feature representation and classification abilities. Moreover, a newly modified network structure (ST-GCN+) is developed to percept the slow actions more accurately, promoting the classification performance for representative actions occurring during CAE analysis. MMST-GCN+ is evaluated on two multimodal human action datasets (MHADs) and one homemade teacher action dataset constructed by investigating the research focus on CAE. It achieves better results of 96.21%, 96.33%, 96.21% on the accuracy, macro F1-score and micro F1-score, respectively, on the Berkeley-MHAD, 91.55%, 91.93%, 91.55% of the corresponding indices on the UTD-MHAD and 79.18%, 80.38%, 79.18% on the same indicators for the self-made dataset Teacher-MHAD.