MT-YOLO: Combination of Multi-Scale Feature Extraction and Transformer in One-Stage Object Detection

Guobin Qi, Xiufeng Zhang, Xingkui Fu · 2023

Object detection is a very important area in computer vision that can be applied to many scenarios and plays a very important role, where YOLO has become the industry standard for effective object detection. In this study, we consider enhancing the capability of YOLO for global feature modeling and design a multi-scale dilated dense connectivity module based on YOLOv5 to extract feature information at different scales and levels to obtain spatially dependent information in images, while adding a Convolutional Block Attention Module(CBAM) to learn details of the detected objects. Then, we use three Transformer Encoder-like integrations to the top of the Neck instead of on the backbone network. In this way, more meaningful, fine-grained and feature-enriched mapping relations can be obtained to extract associations about different image regions. The validation results on the COCO 2017 dataset show that MT - YOLO-S is able to achieve 44.1 % AP while being able to reach a detection speed of 115 frames, in addition, our N/M size model achieves 35.2 % % and 49.5% AP performance in the same experimental setting.

Read the paper · More papers on PaperTik