YOLO11-HP: a high-performance YOLOv11 for occluded pedestrian detection
Hengfeng Gong · 2025
To address the challenges of pedestrian detection in densely occluded scenarios—including limited multi-scale feature fusion, severe feature degradation, and slow inference—we propose YOLO11-HP, an enhanced YOLOv11-based framework. The model integrates three key innovations: an efficient spatial pyramid pooling module (SPPFCSP) with cross-stage partial connections to expand the receptive field and fuse multi-scale features; a combination of CBAM attention and a weighted bidirectional feature pyramid (BiFPN) to enhance occlusion handling and detail capture through channel-spatial enhancement and dynamic feature balancing; and a self-attention-driven dynamic detection head that adaptively weights features across scales, spatial locations, and task dimensions to suppress noise and improve localization. Evaluated on CrowdHuman and WiderPerson datasets, YOLO11-HP achieves a 2.4–2.6% improvement in mAP50 and a 3.1% gain in mAP50:95 over the baseline YOLOv11n, while maintaining computational efficiency (6.6 GFLOPs). It significantly reduces false positives and refines bounding box accuracy, validated by ablation studies confirming each module’s contribution. Designed for real-world applications like autonomous driving and surveillance, YOLO11-HP balances robustness and efficiency, with future work focused on multi-modal fusion and lightweight optimization to enhance real-time adaptability.