YOLO-CSPOKM: Efficient Small Object Detector via CSP Omni-Kernel Model

Wu Wei · 2024

In the field of real-time object detection, the YOLO series has become a mainstream approach due to its exceptional performance. However, its performance on small object detection still has room for improvement. Small objects often struggle with limited feature representation in the P3, P4, and P5 detection layers. Traditional methods to address this issue typically add a P2 detection layer to enhance small object detection capabilities, but this often leads to a significant increase in computation and extended post-processing time. Therefore, developing an efficient and effective feature pyramid tailored for small objects has become an urgent problem to solve. This paper proposes a network specifically optimized for small object detection—YOLO-CSPOKM, which significantly enhances the performance of small object detection while also improving the detection of general objects. Based on the original PAFPN structure, we designed a Small Object Enhance Pyramid: the P2 feature layer is processed using SPD-Conv to extract features rich in small object information and then fused with the P3 layer. Subsequently, the CSP (Cross Stage Partial) strategy is employed to split input features along the channel dimension. One portion of the features is passed through the Omni-Kernel module to effectively capture multi-scale features ranging from global to local levels, while the other portion is concatenated with the Omni-Kernel output via skip connections. Furthermore, the P3, P4, and P5 features are passed through the SSFF module, and their output is added to the results from the CSPOKM module. Finally, the combined features are sent to the detection head, achieving a comprehensive enhancement in small object detection. Experiments on the MS COCO dataset demonstrate that compared to the baseline YOLOv8n model, YOLO-CSPOKM improves [email protected]:0.95 to 39.2%, a 2.8% increase, while maintaining a compact model size. When extended to the YOLOv8s model, the enhanced version achieves an [email protected]:0.95 of 46.4%, representing a 1.9% increase compared to YOLOv8s.

Read the paper · More papers on PaperTik