The Evolution of YOLO: from YOLOv1 to YOLOv11 with a Focus on YOLOv7's Innovations in Object Detection
Yuzhao Luo · Theoretical and Natural Science · 2025
Throughout the evolution of You Only Look Once (YOLO) series, staring from base YOLO to latest YOLOv11, each version takes advantages of different techniques and mechanism, incorporating innovations that enhance object detection capabilities by improving both speed and accuracy. From introduction of anchor boxes in YOLOv2 to multi-scale predictions in YOLOv3 and Cross-Stage Partial Networks in YOLOv4, each iteration has brought unique improvements. In YOLOv7, two major advancements, Extended Efficient Layer Aggregation Network and Planned Re-parameterized Convolution, were introduced to address challenges in feature aggregation and parameter utilization, while maintaining optimal gradient flow. Additionally, advanced label assignment strategies, such as lead head guided label assigner and coarse-to-fine label assigner, further improve learning efficiency. These innovations enable YOLOv7 to set new standards in object detection, especially for applications in autonomous driving, video surveillance, medical imaging, and beyond