A human location and action recognition method based on improved Yolov11 model
Shunyi Chen, Yongkang Liu, Hanqing Zhang, Yi Cai · Discover Artificial Intelligence · 2025
Intelligent monitoring systems often struggle with accurate human detection and action recognition in complex environments such as classrooms. To address this, we propose an improved human behavior recognition framework based on a modified YOLOv11 architecture. A key contribution of this study is the creation of the Student Classroom Behavior dataset (SCB-dataset3), a novel benchmark comprising 5686 images and 45,578 annotations across six behavior classes (hand-raising, reading, writing, phone interaction, head-bowing, desk-leaning) and twelve educational stages from preschool to university. Our model integrates the CBAM attention module and a dual classification head to enhance feature representation and enable simultaneous location and action classification. Optimization via quantization and pruning further boosts deployment efficiency. Experimental evaluations show that our model achieves a mean average precision ( $${\text{mAP}}@0.{5}$$ ) of 0.805 for human detection and 0.722 for action recognition, with an inference speed of 58.2 FPS, outperforming the YOLOv8n and baseline YOLOv11n models by 1.4 and 3.6% in $${\text{mAP}}@0.{5}$$ , respectively, while halving the parameter count compared to the dual YOLOv11 models. These results demonstrate superior performance in both accuracy and efficiency, validating the effectiveness of SCB-dataset3 and the proposed architecture for robust classroom behavior analysis.