Human Posture Detection Model Based on Improved YOLO11-Pose
He Fei · 2025
This paper proposes a human posture detection model based on improved YOLO11-Pose. By integrating top-down and bottom-up detection paradigms, it realizes that a single network inference can output multiple people's position coordinates and pose information at the same time, which significantly improves the computational efficiency. An optimization mechanism based on Key Point Similarity (OKS) was introduced into the design of the loss function, and the evaluation index parameters were directly optimized through the end-to-end training strategy to improve the accuracy of key point location. In addition, a cross-attention feature fusion module (CAFMFusion) and a heat diffusion enhanced C3k2_Heat module are proposed to solve the limitations of the traditional YOLO architecture in complex scenes (such as occlusion and illumination change). Experimental results show that the improved model achieves 70.1% mAP50 on the OC_Human dataset and only needs 10MB parameters, which achieves a balance between accuracy and efficiency, and is suitable for high-demand scenarios such as real-time pose estimation.