Multi-Scale Attention and Transformer-Enhanced YOLO Architecture for Robust Traffic Sign Detection in Complex Visual Environments

Nada Farhani, Mahdi Hermassi, Mohamed Ali Hajjaji, Mounir Zrigui · Intelligenza Artificiale · 2025

Traffic sign detection is a fundamental component of intelligent transportation systems, yet remains challenging due to the small size of signs, visual occlusions, and complex environmental conditions. In this paper, we propose a novel YOLO-based architecture enhanced with multi-scale attention and Transformer modules to address these limitations. Specifically, a Convolutional Block Attention Module (CBAM) is employed to refine spatial and channel-wise features, while a C3 Transformer (C3TR) module introduces multi-head self-attention to capture global contextual information. The proposed enhancements significantly improve the model's ability to detect small and visually degraded traffic signs. Evaluated on the German Traffic Sign Detection Benchmark (GTSDB), our model achieves a [email protected] of 96.75%, [email protected]:0.95 of 81.18%, precision of 97.05%, and recall of 95.07%. Compared to YOLOv5 s, this reflects relative gains of +11.2% in [email protected], + 26.6% in [email protected]:0.95, + 1.6% in precision, and +20.0% in recall, with a 41.9% reduction in model size. It also outperforms YOLOv8, YOLOv7-tiny, and Faster R-CNN, particularly for degraded signs. For real-time deployment on embedded systems, the model is optimized using NVIDIA TensorRT. This optimization significantly reduces inference latency and computational load while preserving high detection accuracy, making the model well-suited for ADAS and autonomous driving applications.

Read the paper · More papers on PaperTik