DSM-IDM-YOLO: Depth-wise Separable Module and Inception Depth-wise Module Based YOLO for Pedestrian Detection
Sweta Panigrahi, U. S. N. Raju · International Journal of Artificial Intelligence Tools · 2022
Pedestrian detection is one of the most challenging research areas in computer vision. Compared to traditional hand-crafted methods, convolutional neural networks (CNNs) have superior detection results. The single-stage detection networks, particularly the state-of-the-art You Only Look Once (YOLO) network, have attained a satisfactory performance without compromising the computation speed in object detection. YOLO framework can be leveraged in pedestrian detection as well. In this work, we propose an improved YOLOv2, called DSM-IDM-YOLO. The proposed model uses a modified DarkNet19 integrated with three new modules, two depth-wise separable convolution modules and one inception depth-wise convolution module, leading to a comprehensive feature of an object in the image. The modules are integrated on top of features from multiple levels of the network. The proposed framework is computationally less expensive owing to its convolution design and a moderate number of layers. It aims to improve performance with minimal computational overhead. The proposed method is compared with state-of-the-art detection methods, i.e., Faster R-CNN, YOLOv2, YOLOv3, YOLOv4-tiny and Single Shot Multibox Detector (SSD). The performance results attest that the proposed method has effectively improved the detection. Three benchmark pedestrian datasets are used for experimental analysis: INRIA, PASCAL VOC 2012 and Caltech.