Head Pose Estimation In Classroom Scenes
Jiayi Shen, Xue Qin, Zhenyu Zhou · 2022
In this paper, we propose an image classification method to recognize head pose in surveillance video of classroom scenes, which is improved on EfficientNetV2-S, called EfficientNetV2-SC. Firstly, the SENet attention mechanism in MBConv module is replaced by the ECA-Net, which can preserve the original channel features in classroom video. Then, a module combining residual structure, channel blending and convolutional attention is proposed to be added between the feature extraction network and the fully connected layer to enhance feature extraction ability and semantic association between channels. Finally, cross-entropy loss function is combined with label smoothing, which can prevent overfitting and improve generalization ability. In the homemade dataset, compared with the original EfficientNetV2-S, the accuracy of our method is improved by 2.7% which is up to 94.8%, and the number of model parameters decrease by 3.55M. In the POSE subset of CAS-PEAL-R1, EfficientNetV2-SC also achieves a significant advantage in experimental accuracy.