Embedding Coordinates in Transformer for Behavior Recognition in Classroom Scenarios

Hongye Zhu, Jinhua Zhao · 2023

The integration of computer vision research into educational contexts has recently attracted significant attention among scholars. One critical downstream application, classroom behavior recognition, poses unique challenges such as: 1) the nuanced differences between behavioral features and 2) the complexity of accurately locating students due to visual obstructions. To tackle these challenges, this study introduces a novel visual Transformer module specifically engineered for behavior recognition in classroom scenarios. Specifically, we first include the Contextual Transformer (CoT) block, which exploits inter-key contextual information to guide the adaptive learning of attention matrices, thereby enabling the extraction of richer behavioral features from classroom images. Then, relative position encoding is incorporated into the CoT block, forming a Position-CoT block (PCoT). The PCoT outputs a fusion of both static and dynamic context representations, and the inclusion of relative position encoding enhances the ability to accurately localize students within the visual frame. Finally, we replace all the standard 3 $\times$ 3 convolutions in ResNet with PCoT to construct a Transformer-style backbone network called PCoTNet. Through experiments, we validate the feasibility of using PCoTNet for classroom behavior recognition. Experimental validation confirms the effectiveness of the PCoT blocks and PCoTNet in classroom behavior recognition.

Read the paper · More papers on PaperTik